Frontier language models sometimes recognize that they are under evaluation and adjust their behavior which can undermine validity of benchmark results.
Yet the field studies it without a shared foundation, conflating flaws of the evaluation with capabilities of the model, and detection with behavioral response.
We ground evaluation awareness in social psychology, decomposing it into an environment component and a model component that separates recognition from propensity.
We operationalize the environment component through eight categorized trigger factors, such as placeholder entities and grading-style output formats, and study recognition and behavior through chain-of-thought monitoring.
Across nine frontier models and four benchmarks, recognition rates depend on the specific pairing of model and benchmark.
Recognition rarely associates with behavioral change, and when it does, the direction depends on the type of evaluation perceived.
Models are also more sensitive to safety than capability evaluations, placing safety benchmark validity at greater risk.
To study which factors each model is sensitive to and how they interact, we propose , a factor-controlled benchmark of 100 paired safety-capability tasks where each of the eight factors can be independently toggled, varying evaluative signals while holding the underlying request fixed.
Through , we find that no single factor uniformly affects all models, but stacking factors progressively raises evaluation awareness across all of them.
Our framework and provide the tools to measure, attribute, and mitigate evaluation awareness, building the foundation for future solutions.