首页 > AI前沿 > TrajLong: Co-Designing Agentic and Long-Context Supervision for Mid-Training

TrajLong: Co-Designing Agentic and Long-Context Supervision for Mid-Training

arXiv自然语言 2026-10-04 13:42 4 阅读 查看原文

LLM agents for coding, search, and workplace tasks increasingly rely on long-context capabilities to effectively aggregate and reason over extended interaction histories.

Recent work has incorporated agent trajectories into mid-training stage, drawing on their naturally long and interaction-rich structure.

Yet how to organize the information within these trajectories into effective mid-training supervision remains underexplored.

In this work

We investigate the relationship between long-context and agent atomic capabilities and introduce TrajLong, a novel framework that compiles trajectories into long-context training tasks with dense supervision, targeting three representative atomic capabilities: evidence grounding, cross-evidence aggregation, and temporal state maintenance.

We mid-train Qwen3-14B-Base and Qwen3-30B-A3B-Base with data compiled by TrajLong, followed by supervised fine-tuning.

Experiments

Experiments on 6 long-context and 12 agent benchmarks demonstrate broad performance gains, with controlled ablations showing improvements over raw and masked trajectory baselines.

Capability-level analyses further reveal task-dependent associations between long-context and agent atomic capabilities.

Findings

These findings suggest that the shared capability demands of long-context reasoning and agent execution provide a principled basis for designing mid-training data to develop downstream agent capabilities.