首页 > AI前沿 > No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays

No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays

arXiv自然语言 2026-10-06 13:00 5 阅读 查看原文

Natural language processing (NLP) models can detect depression-related language in text written near the time symptoms are measured, but whether pretrained transformers can predict depressive symptoms from text written twelve years earlier is largely untested.

In the National Child Development Study, a British birth cohort, we predict probable depressive symptoms at age 23 from essays the same people wrote at age 11.

Our baseline, a logistic regression on six childhood covariates, outperforms every text model that sees only the essay: seven fine-tuned transformers, a bag-of-words model, frozen embeddings and four zero-shot large language models.

Its area under the receiver operating characteristic curve (AUC-ROC) is 0.737 against 0.670 for the best transformer on the primary seed, and no added text score detectably raises the baseline's AUC-ROC.

None of the five domain-pretrained transformers detectably beats its general-domain control after Bonferroni correction.

For long-horizon prediction, the baseline remains the model to beat.