首页 > AI前沿 > Constraint-Aware Training

Constraint-Aware Training

arXiv机器学习 2026-10-02 14:59 11 阅读 查看原文

When generating programs with language models, constrained decoding can apply program analyses to exclude tokens that violate syntax, scope, or typing rules.

However, there is a duplication: standard training already teaches the model to suppress the tokens rejected by these analyses.

This duplication leads to the question: if we will perform some analysis to filter a set tokens out during inference anyways, can we avoid teaching the model the said analysis altogether during training, and does this externalization lead to more efficient models?

This paper defines a general constraint-aware objective satisfying this externalization desideratum and formalizes the benefits of externalization into three concrete theorems about model size and data efficiency.

We show, through a controlled synthetic experiment, that the theorems survive training dynamics:

constraint-aware training yields lower prediction loss at a matched parameter count and data compared to ordinary cross-entropy training, motivating training objectives that incorporate the analyses used during generation.