首页 > AI前沿 > m-Set Adversarial Bandits with Winner Feedback

m-Set Adversarial Bandits with Winner Feedback

arXiv机器学习 2026-10-07 22:10 5 阅读 查看原文

We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations).

By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates.

Our main technical contributions are the information-theoretic lower bounds on the regret.

Experiments on synthetic data confirm our theoretical analyses.