首页 > AI前沿 > Gender bias across LLMs is common and highly heterogeneous

Gender bias across LLMs is common and highly heterogeneous

arXiv自然语言 2026-09-30 01:11 5 阅读 查看原文

Introduction

Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences.

Research Gap

Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs.

Study Approach

We address this gap across ten models released between April 2025 and June 2026, spanning nine vendors, using two paradigms: gender attribution to stereotyped phrases (Study 1) and moral judgment of abuse or torture against a woman or a man to prevent a catastrophic outcome (Study 2).

Study 1: Gender Attribution to Stereotyped Phrases

In Study 1, two of ten models attributed masculine-stereotyped phrases to female writers more often than the reverse, while three models showed the opposite pattern.

Study 2: Moral Judgment of Abuse or Torture

In Study 2, several models converged on a male-disadvantaging asymmetry that was directionally consistent with a documented human tendency to protect female targets from harm, though the specific conditions under which this asymmetry emerged varied by model; three other models, by contrast, showed no variation across conditions.

Conclusions

These results indicate that gender-related biases are common in LLMs. Their direction and magnitude, however, are highly heterogeneous, to the point that some models behave in diametrically opposite ways to others. Bias auditing should therefore be treated as an ongoing, multi-vendor process, rather than a one-time assessment.