首页 > AI前沿 > How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing

arXiv自然语言 2026-10-06 23:32 4 阅读 查看原文

Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data.

We propose reading every probe score against two reference points:

a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs.

We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it.

We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known.

Under distribution shift, probe scores fall and prediction error rises $12$--$15\times$, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation.

We then analyze the real models:

The single-cell foundation model scGPT encodes biological variability only partially.

We also revisit four influential LLM probing studies:

which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users.

Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.