Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it.
When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself.
We test this idea using a novel, forward-only proxy motivated by the norm of the Hessian, alongside common Hessian measures, across molecular property, protein fitness, and drug-response tasks.
Contrary to our hypothesis, geometry does not become more useful as validation Spearman correlation deteriorates: augmenting validation helps some shifts but significantly harms others.
More surprisingly, the proxy still correlates with generalisation gap on most tasks even when Hessian trace and top-eigenvalue relationships are weak or reversed, yet this signal does not reliably identify the deployment-best model.
A curvature bound need not preserve cross-model rankings, and low geometric scores can even favour collapsed predictors.
Thus, a generalisation signal need not be a model-selection signal.