首页 > AI前沿 > C-Instrument: Automating RL Data Generation and Hillclimbing with a Constitution-Grid Instrument

C-Instrument: Automating RL Data Generation and Hillclimbing with a Constitution-Grid Instrument

arXiv自然语言 2026-08-01 02:05 6 阅读 查看原文

Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard.

Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts.

Our finding is that over-refusal improves 22.4% to 12.8%, while under-refusal on adversarial attacks silently worsens 0.27 to 0.33.

C-Instrument

We present C-Instrument, a constitution-grid data instrument that generates the RL training data.

C-LIM

C-LIM, a per-cell learnability score that decides each cell's move: prune, densify, amend, expand.

C-LIM flags the dead-weight data region before any training budget is spent: 187 untargeted rows had bought zero gain, and our method lifts the same region's learning impact 0.733 to 0.80.

Code and the constitution are open-sourced.