首页 > AI前沿 > Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

arXiv自然语言 2026-09-10 04:17 5 阅读 查看原文

Production text-to-SQL pipelines often end with an LLM-as-judge whose agreement with human annotators has never actually been measured.

When we checked ours, the deployed gpt-4o-mini judge agreed with two-author gold at only Cohen's kappa = 0.04 on a disagreement-enriched set and 0.42 on a uniform-random spot-check, over-flagging 77.1% of the human-FAITHFUL cases in the enriched set.

Most of its over-flags trace back to a single mechanism we call GRADE-HALLUCINATION.

A self-hosted Qwen3.6-27B replacement (kappa = 0.72) lands in the same range as Claude Opus 4.7 (kappa = 0.71); the head-to-head is underpowered at n = 96, but for the deployment decision that hardly matters, since Qwen costs roughly 1/300 as much per call.

Ensembling does not help for free.

Pairing the weak judge with a stronger one degrades agreement, whereas three strong judges under unanimity routing reach kappa = 0.79 at 89.7% auto-coverage.

Applied out-of-domain, the same audit recipe flags 25.5% of BIRD-financial's expert-authored gold SQLs as candidate gold-SQL issues under our annotation protocol.

Code and pre-registration are at https://github.com/JamesL404/synca-audit.