Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely signaling task completion.
Existing monitoring approaches mainly detect deception after it appears in observable actions or outputs.
Investigation on Predicting Deceptive Behavior
In this paper, we investigate whether deceptive behavior can be predicted from an agent's internal representations before it becomes externally visible.
We frame deception monitoring as a trajectory-level representation analysis problem and align agent trajectories around key decision points.
Using hidden states extracted before these points, we show that future honest and deceptive outcomes can be reliably distinguished, with predictive signals remaining detectable several model calls before the final decision.
Temporal Evolution of Signals
We further characterize the temporal evolution of these signals: deception-related representations are weak early in execution but become increasingly identifiable as trajectories progress, while transferable structure can emerge before the strongest decision-adjacent signals appear.
Intervention and Findings
Finally, we intervene on the identified honest-deceptive representation directions during inference and find that activation steering reduces downstream deceptive behavior, suggesting that these representations influence agent decisions.
Our findings indicate that agent deception is an evolving internal process that can be detected and potentially mitigated before it is expressed externally.