As people turn to LLMs for social advice, understanding their behavior in such contexts becomes essential.
In this work, we focus on behavioral dispositions: the underlying tendencies that shape responses in social contexts.
We introduce STAR
STAR builds on established psychological questionnaires, adapting their items into realistic advice-seeking scenarios, as self-report may not transfer to actual advisory behavior.
Dataset Construction
Using STAR, we construct a dataset of 23k scenarios, each validated by 3 raters and annotated with preferences from 10 participants.
Findings Across 25 LLMs
- when human consensus is high, frontier models can fail to reflect it in 15-20% of cases, and smaller models fail at substantially higher rates;
- when humans disagree, LLM recommendations are substantially less diverse than human choices, both within individual models and even across models from different providers, potentially narrowing the range of options users are guided toward;
- LLMs' self-reported values are poor predictors of their recommendations.
Supporting Future Research
To support future research we make our dataset and code publicly available.