首页 > AI前沿 > DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks

DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks

arXiv机器学习 2026-10-08 09:53 4 阅读 查看原文

Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data.

However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance.

This paper

from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:

  • (i) ineffective utilization of prompts length under diverse task complexities,
  • (ii) semantic irrelevance of randomly sampled prompt segments,
  • (iii) misleading supervision induced by fragmented and discontinuous trajectories.

To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality,

featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC).

Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks.

Experimental results

on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods,

achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets,

demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.