Multimodal meme understanding is increasingly used to analyze socially sensitive content, yet existing models often exhibit biased behavior when interpreting economic dependence and social roles under ambiguity.
Many memes express economic relationships through sparse text or symbolic visual cues, providing insufficient evidence for gendered attribution.
In such underspecified settings, models tend to rely on pretraining correlations, leading to hallucinated and stereotypical economic role assignments.
In this work
We study gendered economic dependence in image-text memes through the lens of contextual sufficiency and identify epistemic overcommitment-inferring roles without adequate evidence-as a primary source of bias.
Proposed Framework
We propose CGER-Net, a context-grounded multimodal framework that estimates whether the input provides sufficient evidence for gendered economic reasoning and applies evidence-gated inference to enable confident attribution when cues are explicit while favoring principled abstention otherwise.
Evaluation
We evaluate CGER-Net on EconMeme-GE, a curated dataset of image-text memes annotated as Men, Women, Neutral, or Ambiguous.
Across strong contemporary multimodal baselines, CGER-Net reduces Gender Overcommitment Rate by up to 44% on ambiguous instances while maintaining comparable accuracy on unambiguous cases.
Human Evaluation
Human evaluation further shows that 79% of generated rationales are judged as epistemically aligned with the available evidence.
Conclusion
These results highlight the importance of modeling when not to infer for reliable and responsible multimodal analysis.