Depression has no diagnostic blood test.
Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals?
We pre-registered 880 language-model raters
Crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire.
Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%.
Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average.
Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five.
A locked analysis of 86 new interviews
Reproduced the main pre-registered findings.
Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently.
Calibration repaired much of the rater dependence without securing agreement about individuals.