首页 > AI前沿 > When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction

When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction

arXiv自然语言 2026-09-24 05:05 7 阅读 查看原文

Intrinsic self-correction asks a language model to revise its own answer without receiving new external evidence.

A second pass can recover mistakes, but it can also overturn answers that were already correct.

We study this trade-off across 29 open-weight LLMs on BoolQ, GSM8K, and Corr2Cause by tracking correctness transitions between initial and revised answers.

Aggregate accuracy can conceal substantially different revision behavior: for example, Llama-3.1-8B improves by 25.5 percentage points on GSM8K, while refinement changes 19.1% of initially correct answers into wrong ones.

A controlled BoolQ study further shows that refinement prompts shift the balance between recovery and harm.

We then compare three runtime choices:

  • keeping the initial answer
  • always accepting the revision
  • selectively invoking revision using signals available after the initial response

The comparison identifies settings where learned gating is useful and others where a simpler unconditional policy performs better.

These results suggest treating intrinsic self-correction as a revision policy rather than as a uniformly beneficial second pass, and evaluating it through both the corrections it recovers and the errors it introduces.