Visual entailment (VE) asks whether an image supports, contradicts, or leaves undecided a textual hypothesis.
Strong results come from fine-tuning large vision-language models on labelled data, while zero-shot and hybrid approaches remain far behind.
A VE hypothesis often bundles several visual claims, yet existing zero-shot methods reason over it as a single unit.
We propose Atomic Visual Entailment (AVE)
We propose Atomic Visual Entailment (AVE), which decomposes the hypothesis into atomic facts, produces candidate predictions from both the full hypothesis and its facts using frozen vision-language models, and predicts the final label with a lightweight classifier trained only on how those candidates behave.
We find that decomposition helps only when the hypothesis context is preserved: judging facts in isolation is worse than not decomposing at all.
Full-hypothesis and atomic prediction make complementary errors, and learning which to trust recovers far more of that complementarity than majority voting, reaching 0.803 test accuracy on SNLI-VE without fine-tuning any vision-language model.
AVE also localises the visual evidence behind its prediction without region-level supervision
These results suggest that learning which candidate prediction to trust can close much of the gap to fine-tuned systems, offering a practical alternative where fine-tuning a vision-language model directly would need more labelled data or compute than is available.