Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks.
A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from.
We study how such adaptivity is learned in a controlled location-estimation problem.
Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$.
We also provide the example of a symmetric Gaussian mixture, for which a rate of $σ^2_n/n$ can be attained.
On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function.
A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range.
We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow.
With $\widetildeΩ(n^{1+ε})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^ε$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime.
These guarantees extend to new locations and longer contexts.
A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast.
Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection.
End-to-end experiments recover the predicted specialization.