首页 > AI前沿 > The Missing Minimal Pair: Stereotype Evaluation in LLMs

The Missing Minimal Pair: Stereotype Evaluation in LLMs

arXiv自然语言 2026-10-07 01:41 5 阅读 查看原文

A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences.

We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences.

To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation.

First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes.

We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes.

Second, we introduce two evaluation metrics tailored to the dual minimal pair setup.

One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes.

This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models.

Our code is available at https://github.com/stepanat/missing-minimal-pair/.