Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution.
While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Bayesian behavior, standard supervised fine-tuning (SFT) on the true answers for the task falls short of it.
But behavior alone does not tell us $\textit{why}$ tuning on a $\textit{Bayesian}$ or an $\textit{oracle}$ (true answers) signal differs: whether the resulting LM represents Bayesian beliefs, acts on them, or turns them into a choice the way Bayes' rule does.
To compare them, we formulate increasingly demanding requirements for an LM to count as a Bayesian decision maker, spanning its behavior, representations, and computations, and test them on a flight recommendation task.
The Bayes-trained LM acts Bayesian, encodes quantities of Bayes' rule in its middle layers, and uses the encoded belief for the recommendation to a certain extent.
The oracle-trained LM differs from it both in the beliefs it holds and whether it reads beliefs out into recommendations.
Exchanging beliefs between the LMs transfers a part of the Bayesian advantage.
Bayes fine-tuning thus installs usable Bayesian beliefs in an LM for reasoning under uncertainty in a way standard SFT on oracle answers cannot, highlighting the advantage of nuanced supervision.