首页 > AI前沿 > When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions

When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions

arXiv机器学习 2026-09-30 04:57 7 阅读 查看原文

A valid curvature upper bound need not justify either a robustness certificate or an intervention on an intrinsic predictor property.

We demonstrate this distinction for a last-layer relative-flatness proxy used in both settings.

First

First, empirical-risk stationarity does not eliminate pointwise first-order loss terms:

at a finite global empirical-risk minimum, the retained certificate expression underestimates a loss increase by over $210\times$.

We derive a globally valid, gauge-invariant feature-space repair.

Second

Second, common-row softmax shifts preserve predictions and the exact contraction while making the proxy unbounded.

Even standard reference-class choices double it on average relative to the centered representation.

For a single fixed-feature example with at least three classes, scalar retuning generically cannot align the induced probability updates.

Row centering gives the orbit-minimized bound and restores value and full-model gradient invariance under this symmetry.

Experiments

Across 45 paired one-step tests on algorithmic and image models, amplified shifts separate raw-regularized predictors while quotient-regularized predictors remain aligned.

Experiments (Continued)

Long-horizon CIFAR-10 experiments show substantial, reversible suppression of generalization, while evidence for selective delay after memorization is less consistent.

Together, these results show that validity as a curvature upper bound does not by itself justify either inversion into a robustness certificate or differentiation into an intrinsic training intervention.