Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers.
This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference.
Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries.
There, symmetry enables a precise description of the spectra, forcing high-dimensional kernels and eigenvalues of large multiplicity.
Returning to the original configuration breaks the Hessian symmetry and thereby produces the observed hierarchy of clusters and outliers.
The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK.
The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.