首页 > AI前沿 > Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses

Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses

arXiv自然语言 2026-10-08 20:03 6 阅读 查看原文

Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch.

BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict.

We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution.

Under tight patch budgets, entropy-triggered layouts skip these positions, and accuracy on them collapses.

In Meta's BLT-1B

In Meta's BLT-1B with patch starts on 10% of bytes, the entropy rule puts a patch start at 16% of the computed results in GSM8K solutions and gets 19.0% of them exactly right;

a boundary after each "=" at the same patch count gets 51.8%, and entropy combined with a label-free boundary-dependence signal gets 67.1% (default layout at 26% of bytes: 76.8%)

Adapting BLT-1B

The gap survives adapting BLT-1B to the budget with low-rank fine-tuning (32.9% vs 72.7%, three runs per rule, paired p < 1e-200) and grows with model size in byte models trained from scratch at a 10% budget:

at 1M, 12M and 50M parameters, boundary dependence beats entropy on final answers by -1.6, +10.1 and +19.8 points, and at 50M it gets 35.9% of computed results against 13.9% (3 seeds each)

Entropy-jump Rule

BLT's entropy-jump rule helps neither target at 50M.

Scratchpad Patching

The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%

Effect on Computed Values

The effect is specific to computed values: copies and lookups gain little, and values the model cannot compute gain nothing.

Boundary Dependence

Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels:

combined with entropy it beats the hand-written rule on computed results.