首页 > AI前沿 > Fully Interpretable Minimal Transformers: From Geometry to Algorithm

Fully Interpretable Minimal Transformers: From Geometry to Algorithm

arXiv机器学习 2026-10-07 19:01 5 阅读 查看原文

We present a framework for building and interpreting minimal transformer models.

By constraining a transformer's embedding dimension and head size to 2, we enable full two-dimensional visualization of its internal representations.

Embeddings, query/key/value transforms, attention outputs, residual streams, and decision boundaries can all be seen directly.

Our central claim is that the learned geometry implies an algorithm; the arrangement of points and boundaries in R^2 can be read as a step-by-step procedure.

We train a transformer on a simple task where it must produce the most recently observed even number whenever the '+' operator appears in a sequence of digits.

Once trained, we visually walk through every step of the transformer's computation.

We show how the model embeds the tokens and their respective positions in the sequence, transforms them via the Q, K, and V matrices, uses the dot product between the Q and K representations to form the attention matrix, and uses the attention matrix to select values that move the representation of each input token to the region of the domain of the output layer that will correctly predict the next token.

We introduce a suite of interpretability visualizations that make the algorithmic interpretation of this procedure explicit.

Our framework offers a pedagogical and experimental testbed to explore how transformers use informational geometry to implement next-token prediction.