首页 > AI前沿 > Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis

arXiv机器学习 2026-10-06 13:57 5 阅读 查看原文

Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.

A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from.

We study how such adaptivity is learned in a controlled location-estimation problem.

Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$.

We also provide the example of a symmetric Gaussian mixture, for which a rate of $σ^2_n/n$ can be attained.

On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function.

A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range.

We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow.

With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime.

These guarantees extend to new locations and longer contexts.

A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast.

Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection.

End-to-end experiments recover the predicted specialization.