首页 > AI前沿 > Language-model ratings of depression reflect the rater more than the patient

Language-model ratings of depression reflect the rater more than the patient

arXiv自然语言 2026-10-06 23:08 5 阅读 查看原文

Depression has no diagnostic blood test.

Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals?

We pre-registered 880 language-model raters

Crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire.

Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%.

Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average.

Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five.

A locked analysis of 86 new interviews

Reproduced the main pre-registered findings.

Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently.

Calibration repaired much of the rater dependence without securing agreement about individuals.