Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data.
However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance.
This paper
from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:
- (i) ineffective utilization of prompts length under diverse task complexities,
- (ii) semantic irrelevance of randomly sampled prompt segments,
- (iii) misleading supervision induced by fragmented and discontinuous trajectories.
To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality,
featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC).
Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks.
Experimental results
on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods,
achieving an average improvement of 11.73% on optimal datasets and an improvement of 13.34% on suboptimal datasets,
demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.