首页 > AI前沿 > What happens when an LLM never sees material beyond fifth grade?

What happens when an LLM never sees material beyond fifth grade?

Hacker News 2026-08-16 15:37 1 阅读 查看原文
LittleCurriculum An 88B-token corpus distilled from FineWeb-Edu through a five-stage filtering pipeline aligned with Common Core standards (K–5). Concepts, facts, and vocabulary taught above Grade 5 are explicitly excluded. LittleLearner Three scales (0.6B / 1.3B / 5B) trained from scratch on LittleCurriculum: chattable models with an interpretable knowledge boundary. Each ships with a matched Unfiltered control for clean comparison. Elicitation, not acquisition In our experiments, scaling, SFT+GRPO post-training, and in-context learning amplify what the curriculum taught, but none meaningfully improves out-of-scope performance, indicating that the pretraining filter sets the effective capability ceiling.