While AI weather models now rival operational forecasts, how they represent the atmosphere internally remains an open question:
feature attribution reveals which input patterns matter, not what the model computes or how it combines information internally.
We train sparse autoencoders (SAEs) on GraphCast to uncover its learned concepts, using atmospheric rivers as our phenomenon of focus.
Both standard and Matryoshka SAEs show GraphCast computes atmospheric river intensity, measured by integrated vapor transport (IVT), as a stable internal variable, despite IVT being neither an input nor a target.
In contrast to the unstructured concept retrieval of the standard SAE, the Matryoshka SAE orders concepts by importance and exposes their relations.
Atmospheric river concepts persist across depth and direct interventions confirm causality.
This method offers a way to find internal variables and determine which of them the model actually relies on, which is a prerequisite for asking whether those variables remain meaningful as the phenomenon changes under a warming climate.