首页 > AI前沿 > Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

Explicit Trajectory Diversity for RL-Based Post-Training of LLM Agents

arXiv机器学习 2026-09-30 10:35 7 阅读 查看原文

LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory.

Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation.

While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task.

In this work

In this work, we study explicit trajectory diversity in RL-based post-training for LLMs.

Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretable behavioral representation, and then measure diversity as a set-level functional over the resulting descriptor matrix.

Building on this formulation

Building on this formulation, we introduce Trajectory-guided Joint Policy Optimization(TJPO), a single-policy framework that optimizes explicit diversity over sampled trajectory groups, avoiding the need for population-based policy training, and instantiate it within group-based policy optimization through trajectory-level learning signals.

This design makes the diversity objective both interpretable and controllable.

Experiments

Experiments on Sokoban and ALFWorld show that TJPO improves task-specific trajectory diversity while maintaining competitive task performance.

Descriptor and trajectory analyses show that the learned variation follows the specified behavioral dimensions and includes distinct successful strategies.

Extra experiment results suggest that explicitly shaping trajectory diversity can help LLM agents satisfy user requirements and remain effective when task conditions change.