首页 > AI前沿 > Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation

Reinforcement Learning with Segment Reward Feedback under Linear Function Approximation

arXiv机器学习 2026-10-06 20:45 6 阅读 查看原文

Classical reinforcement learning (RL) assumes that a reward is observed for every visited state-action pair.

However, in real-world applications such as autonomous driving, such fine-grained feedback can be costly or difficult to collect, whereas trajectory-level feedback may be too sparse for efficient learning.

To provide a general feedback model bridging these two extremes and handle large state spaces, we study RL with segment reward feedback under linear function approximation.

Our work answers how the granularity of segment feedback and the choice of segmentation influence learning.

For equal-length segments with known transitions, we design algorithms $\bitssegd$ and $\edlinucbsegd$ for binary and sum feedback types, respectively.

Nearly matching lower bounds are established.

For equal-length segments with unknown transitions, we develop a unified $\seglsvits$ framework with two instantiations for binary and sum feedback, which carefully integrates the posterior estimated reward parameters into least-squares value iteration.

Finally, to investigate whether segmenting according to state-action features can further expedite learning, we design an algorithm $\uneqsegbitsd$ that allows arbitrary segmentations.

The resulting regret bound shows that under the usual elliptical potential analysis, the influence of state-action features on the regret appears only through logarithmic factors, and equal segmentation achieves the best performance.