首页 > AI前沿 > When Rank Rises as LLMs Degrade

When Rank Rises as LLMs Degrade

arXiv机器学习 2026-10-07 16:22 4 阅读 查看原文

Post-training adapts language models in non-stationary environments.

Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade.

We show that this assumption is unsafe for LLM post-training.

In a controlled study of Qwen3-0.6B

with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations.

Covariance effective rank rises to nearly twice its healthy value.

This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest.

By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds.

Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone.

We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues.

On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range.

In a pre-registered shared-prefix

leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction.

However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed.

Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.