Alignment does not eliminate behavioral errors in language models.
Models may still refuse benign requests, call unnecessary tools, or yield to false user claims.
Current methods mitigate such errors as a computation problem, and rarely explore if the desired behavior is already encoded in the model's representation.
Motivated by the observation that behavior-relevant information remains linearly decodable from the final hidden state even when the resulting logits produce the undesired behavior, we introduce HeadEdit, a gradient-free method that calibrates model behavior through the unembedding matrix.
HeadEdit extracts a low-rank behavioral subspace from paired completions and uses each prompt's coordinates within it to generate a vocabulary-wide correction, thereby implementing implicitly adaptive steering without manually specified target tokens or parameter updates.
Experimental Results
HeadEdit improves all nine experimental settings across three tasks and three model families, with negligible inference overhead and no systematic loss of general capabilities.
It also reveals a connection to gradient-based alignment.
HeadEdit's low-dimensional representation partly predicts how preference tuning changes output logits on unseen prompts.
The subspace learned from the model can also be reused after tuning, improving performance without re-extracting or retuning.
Conclusion
These results show that HeadEdit provides a practical, lightweight, and interpretable way to calibrate model behavior through the unembedding matrix.