Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored.
This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions.
Methods
Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing.
Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels.
Results
Results revealed systematic and condition-consistent response modulation.
Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models.
Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity.
Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios.
An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone.
Discussion
The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited.
More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.