LLM对非英语母语用户的挑战
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers.
非母语英语用户响应质量差距的原因
Which specific features of non-native English drive this gap remains unclear, because fluency is itself a composite of mechanical accuracy, vocabulary use, organization, and discourse coherence.
FABLE数据集介绍
Here, we introduce FABLE, a controlled dataset of 190,911 English prompt variants derived from 174K real user prompts for writing-related tasks.
评估结果
Evaluating responses from 34 open-weight LLMs, we find a clear asymmetry; while models do not propagate surface errors such as misspellings into their outputs, models do mirror higher-level rhetorical and lexical qualities present in the user's prompt.
Further, the overall quality of responses differs substantially between the least- and most-fluent prompts.
关键发现
These results highlight a key LLM performance disparity for non-native English LLM users, resulting in both lower-quality and less-fluent answers.