首页 > AI前沿 > PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

arXiv自然语言 2026-10-03 07:07 6 阅读 查看原文

Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses.

A socially competent agent must infer users' unstated goals, respect their preferences, and adapt as the conversation unfolds.

These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable.

To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback.

Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population.

Empirical evidence shows that PB-GRPO improves models' social behavior over strong reinforcement learning baselines in our simulated environment.