Test-time compute has become a central way to improve code generation: systems sample multiple candidate programs and use verifier-visible evidence to select the final output.
This paradigm implicitly assumes that the verifier provides a corrective signal independent from the generator.
We challenge this assumption under misleading task premises.
When the generator and verifier share a false premise, they become coupled through a mistaken belief: the generator produces premise-consistent shortcuts, while the verifier supplies evidence that fails to expose them.
Consequently, the selector may choose a hidden-test-wrong candidate even when a hidden-test-correct program exists in the pool.
We call this failure mode Verification Trap.
Experiments and Results
Across three code-generation benchmarks and five code models, false premises consistently degrade first-sample correctness, reduce selector-chosen correctness after 64-sample test-time selection, and amplify recoverable mis-selection.
Mechanistic Insights
Mechanistically, verifier-written tests inherit the premise-level blind spot, reshaping verifier-visible candidate space away from hidden-test correctness.
These traces make Verification Trap predictable before hidden execution: a lightweight gold-free predictor using verifier-visible features reaches 0.846 AUROC.
Identification of Mitigation Axes
Our results identify decoupled evidence as a key mitigation axis: coupled scaling provides limited recovery, whereas premise-agnostic robustness auditors recover substantial oracle headroom.