首页 > AI前沿 > A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese

A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese

arXiv自然语言 2026-10-04 11:19 5 阅读 查看原文

This study analyzes the predictive power of LM-derived, token-level surprisal on Mandarin Chinese reading times.

We first propose the Shortest Matching Sequence (SMS), an alignment scheme that maps between the word segmentation assumed by eye-tracking corpora and the LMs' subword tokenization, as the two tokenizations often disagree in the context of Mandarin Chinese.

Then, using a suite of Chinese-Pythia models (14M-1.4B) trained on scratch with 30B tokens, we examine how well surprisal predicts first fixation duration, gaze duration, and total reading time in three paragraph-level eye-tracking corpora of Mandarin Chinese (GECO-CN, HKP, and MECO).

Contrary to previous null findings, our results show that surprisal is predictive of Chinese reading times.

However, whether predictive power scales with model size and the amount of training is corpus-specific: bigger models predict better in GECO-CN, whereas inverse scaling emerges in HKP and, at the largest sizes, in MECO.

Subsequently, we tested one possible explanation for the inverse scaling in HKP and found that checkpoints whose surprisal remains closer to $n$-gram statistics are better predictors of reading.

All in all, the predictive power of surprisal on Chinese reading time measurements is corpus-specific, which cautions against drawing scaling conclusions from a single corpus.