首页 > AI前沿 > CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning

arXiv机器学习 2026-08-18 11:29 6 阅读 查看原文

Multi-step causal reasoning requires chaining inferences where each step constrains the next. An early error propagates silently, and a correct answer reached via flawed logic evades outcome-level detection.

In specialized domains, teacher LLMs err on intermediate steps, safety constraints restrict cloud distillation, and shifting conditions demand adaptation, leaving self-evolution as the practical route.

Naive self-evolution can collapse: outcome-only rewards let the model exploit distributional shortcuts, and weak self-evaluation reinforces spurious paths into stable failure patterns.

We exploit a key asymmetry: generating a correct chain is hard, but verifying a single step is easy.

Many high-stakes domains admit a deterministic, queryable oracle, a physics simulator or rule engine over codified constraints. It checks asserted steps without teacher-level ability and abstains beyond its rules; it can check what the model asserts, never replace it.

This enables CoEvo, an oracle-grounded self-evolution framework where a single model alternates between Proposer and Solver.

As Solver, the model generates competing chains; intra-group debate exposes disagreement steps, a proxy for the capability boundary, and the oracle adjudicates them into process-level supervision.

As Proposer, the same model constructs progressively harder scenarios inside oracle constraints, steering the curriculum toward deep multi-hop chains.

Both roles are updated jointly, so training pressure co-evolves with the model.

On industrial, clinical, and legal multi-step causal reasoning benchmarks, CoEvo enables an 8B LLM to sustain self-evolution, surpassing distillation baselines and the strongest proprietary reference on path correctness (82.1% vs. 71.4%)

The trained model generalizes to unseen categories and systems, preserving root-cause accuracy.