首页 > AI前沿 > Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

Loop-Free Inverse Reinforcement Learning via Sequential Value Recovery with Q-Score Matching

arXiv机器学习 2026-09-30 12:21 7 阅读 查看原文

Inverse Reinforcement Learning (IRL) aims to recover a reward function that explains expert demonstrations.

Existing IRL methods typically rely on a bi-level optimization procedure that alternates between reward learning and policy optimization, leading to substantial computational burden and training instability.

In this work, we introduce a different route that eliminates policy optimization entirely by leveraging diffusion policies.

Our key insight is that a diffusion policy encodes the action-gradient structure of the optimal soft Q function, enabling reward learning to be cast as a sequence of value recovery problems, thereby allowing us to bypass reward-policy loops inherent in prior IRL methods.

Specifically, our method proceeds in three stages:

  • (I) recovering the optimal soft Q function via action-gradient matching and estimating the corresponding soft value function (LogSumExp of Q values) in a way inspired by Gumbel regression;
  • (II) calibrating these soft values by inferring a state-dependent offset;
  • (III) extracting the reward by enforcing Bellman consistency.

This leads to Loop-Free Inverse Reinforcement Learning (LFIRL), a fully offline algorithm that operates in a simple, loop-free, and sequential manner.

LFIRL is simple to implement and significantly improves training efficiency while maintaining strong reward recovery performance.

Empirically, across Maze, Franka Kitchen, Adroit Hand Pen, and Push-T benchmarks, LFIRL achieves 2-3x speedup over the fastest baselines, while matching or surpassing state-of-the-art methods in reward recovery quality.