首页 > AI前沿 > Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning

arXiv自然语言 2026-09-29 23:44 5 阅读 查看原文

Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization.

The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds.

In a controlled experiment

a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36.

Intervention Score

We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and use it to form the static intervention-value baseline (Static-IV).

Selective dynamic intervention re-ranking (DIR-R)

We then introduce selective dynamic intervention re-ranking (DIR-R), which revisits that subset only when a calibrated probe justifies the comparison.

On the Natural-TOFU dataset

our method has positive descriptive margins in 19/20 comparisons between methods and objectives, although several are near zero.

On the LACUNA localization-precision benchmark

our mean terminal utility is higher in all six negative preference optimization (NPO) and SimNPO comparisons:

NPO margins range from +0.431 to +0.848, and SimNPO margins range from +0.503 to +0.571.

The gradient-difference (GradDiff) objective

The gradient-difference (GradDiff) objective reveals substantial field dependence.

Relative to Static-IV, the primary four-field GradDiff evaluation has six wins, six ties, and no losses, with mean and median paired gains of +0.165 and +0.0025.

The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision

The evidence supports separating localization, initial intervention selection, and checkpoint-dependent support revision.