Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting.
Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good.
We seek to understand when such forgetting is not catastrophic.
A minimal associative memory reproduces these dynamics with three ingredients:
- keys with shared structure
- concentrated new values
- normalization in the network
Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion.
Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model.
Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic.
Which one dominates depends on whether the new data move old memories together or apart.