首页 > AI前沿 > Who Warmed the Archives? LLMs Overestimate Historical Warmth

Who Warmed the Archives? LLMs Overestimate Historical Warmth

arXiv自然语言 2026-09-28 01:15 5 阅读 查看原文

Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand.

Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias.

Methodology

Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016).

  • All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p<0.01).
  • The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models.
  • The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error.

An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era.

Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.