Fluent LLM explanations may not follow the evidence from a structured system.
We present VERITYGATE, a four-gate checker for declared evidence IDs, entities, numbers, and claim types.
It checks a fixed schema; it does not verify every fact in the prose.
At r=0 and r=1, we test 900 instances per setting (450 grounded-ungrounded pairs) with GPT-4o-mini, Llama-3.3-70B, and Claude Sonnet 4.6.
Under this schema-level contract and before repair, 80.3% of mini claims and 47.9% of Sonnet claims fail.
These are verifier rejection rates, not prose-hallucination rates.
One repair pass raises claim survival from 19.7% to 28.0% for mini and from 52.1% to 54.3% for Sonnet.
Verified claims per example change by +0.14 for mini, -0.71 for Llama, and -0.47 for Sonnet, so survival and output volume must be reported together.
A second Sonnet pass gives no clear gain.
At r=1, Gate 4 covers 97.0%, 98.7%, and 100% of failing claims for mini, Llama, and Sonnet.
Small human studies support the rules but show gaps between schema checks and correct prose.
A domain-specific GPT-4o judge test shows an order effect, so it is only a usefulness check.
We release the code and data.