Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans.
Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-making (Cao, et al., 2019).
In Experiments
In 5 experiments with 48 experimental conditions employing over 5,000 trials, GPT-4o and Claude 3.7 Sonnet were tested on two variations of a Bayesian reasoning task.
We also assessed LLM evaluation of the competence and morality of a hypothetical person who had offered the same reasoning task as them.
LLMs hovered near human performance on the Bayesian task, though their reasoning was more rule-based and rigid.
Surprisingly, like humans but to a greater extent, LLMs also demonstrated the same hypocrisy in condemning others who, like them, had deployed Bayes' rule.
In demonstrating Bayesian hypocrisy, LLMs highlight a humanlike error of a dissociation between self-performance and other-judgment, and caution against their use in domains where statistical fidelity and fairness norms collide.