A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans.
This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale.
We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs.
Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor.
Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks.
Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods.
A robust pattern across different modes of bias is that:
- A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions,
- Content-similar benchmarks do not necessarily cluster together,
- The $g$ factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks.
Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development.
This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.