首页 > AI前沿 > When Accuracy Gaps Fail to Certify: Auditing Cross-Domain Recalibration of LLM Judges

When Accuracy Gaps Fail to Certify: Auditing Cross-Domain Recalibration of LLM Judges

arXiv机器学习 2026-08-23 14:51 6 阅读 查看原文

A scalar recalibration map fitted for an LLM judge on one task can fail when the task distribution changes, but the source-target accuracy gap is often treated as a proxy for that failure.

We test what this gap can predict and what it can certify across thirteen judges, two generators, eight domains, and 1,176 predeclared transfers.

After accounting for mean score shift, the gap yields a population lower bound on target calibration error, yet identical gaps can induce opposite transfer outcomes.

Exact importance weighting recovers target proper loss under covariate shift, so failure of an estimated weighting pipeline does not by itself establish conditional shift.

A finite-sample simultaneous lower certificate converts the population bound into a one-sided rejection rule using audit labels disjoint from evaluation outcomes.

The leak-free gap correlation is 0.25 (95% CI [-0.09, 0.55]), falls to 0.09 on the second generator, and does not support a generator-invariant association.

The certificate retains nominal coverage but has power 0.13 even at m=1024, whereas target-domain temperature scaling with 16 labels reaches harm rate 0.09, compared with 0.34 for source-fitted Platt scaling.

Accuracy gaps are therefore weak warning signals for scalar probability transfer, not deployment certificates.