首页 > AI前沿 > Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits

Structure, Not Belief: Correlated Thompson Sampling from LLM-Derived Covariance in Combinatorial Semi-Bandits

arXiv机器学习 2026-10-06 06:35 3 阅读 查看原文

Combinatorial Thompson sampling (CTS) draws independent posterior samples for every arm, so its exploration dynamics ignore any relation among arms.

We study a minimal change to those dynamics: an LLM is queried once for a partition of the arms, the partition becomes a positive-definite correlation matrix $Σ$ through an RBF kernel on cluster ranks, and the per-round posterior sample is drawn with covariance $Σ$ while the Beta posteriors are updated from real rewards only, so the LLM shapes how the sampler moves, not what it believes.

We give a self-contained Bayesian regret bound for the idealized Gaussian sampler whose information gain splits into a $K\log T$ term from the $K$-cluster structure and a ridge term that grows to $d\log T$: the $\sqrt{d/K}$ improvement over independent sampling is a finite-horizon transient, exact only as the within-cluster correlation tends to one.

The correlated sampler reduces regret by 19% over CTS on 16 synthetic Bernoulli families at $T=2{,}500$ (6-7% at $T=25{,}000$ with data-adaptive kernels) and by 41% on the Microsoft MIND-small news benchmark ($d=200$ real articles), while pseudo-observation warm starts give nothing.

An LLM-free ablation with a simulated oracle of controlled quality shows that on unstructured instances the gain is a property of the kernel shape (a random partition, or a plain tempering of the sampling noise, reproduces it), while belief injection at matched oracle quality never helps.