首页 > AI前沿 > Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

arXiv自然语言 2026-10-07 00:57 5 阅读 查看原文

Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role.

We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces.

A signed-number-at-pointer baseline explains most recovery over exact quotation checks.

To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts.

These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld.

Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence.

A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations.

The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed.

For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.