首页 > AI前沿 > Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts

Measuring Cultural Alignment Beyond the Average: A Framework for Evaluating Maternal-Health LLM Interactions in Indian Contexts

arXiv自然语言 2026-10-08 17:35 4 阅读 查看原文

Existing evaluation methods for healthcare LLMs primarily assess factual correctness, safety, and fluency, while providing limited insight into whether generated interactions reflect culturally situated healthcare reasoning.

This limitation is particularly important in maternal health, where care decisions are shaped by social and relational norms.

We introduce MH-INDIC

MH-INDIC, a culturally grounded evaluation framework for maternal-health interactions in urban and semi-urban North Indian contexts that operationalises cultural behaviour through ten dimensions of maternal-health reasoning.

Using a 26-item survey administered to 102 pregnant and postpartum women from urban and semi-urban North India, we evaluate ten LLMs.

We distinguish population level cultural alignment from profile-level behavioural variation.

Although several models approximate the human population-level distribution, all evaluated systems exhibit substantially lower variation across demographic and household profiles than the human cohort, revealing a gap between aggregate alignment and profile-conditioned sensitivity.

Downstream Application of MH-INDIC

As a downstream application of MH-INDIC, we use the strongest-aligned proprietary and open-source models to generate culturally conditioned maternal-health dialogues under zero-shot, self-conditioned, and human-grounded prompting.

Human-grounded conditioning produces stronger profile alignment and dialogue quality ratings, suggesting that measured cultural profiles can improve the cultural grounding of generated interactions