A Strassen-type algorithm has many realizations with the same exact product and multiplication count yet different fp8 error because basis changes reshape coefficient geometry, posing the question of which to run.
No current account settles this: classical stability controls worst-case $\ell_1$ growth, not the expected-error magnitude, and the Dumas--Pernet--Sedoglavic optimizer could only be called probably optimal, its global optimality unproved.
To settle this, we attach to each realization a coefficient functional $Φ$, a scalar summary of its coefficient geometry, which we minimize over the change-of-basis orbit.
This Kempf--Ness problem on a Hadamard manifold lets us certify the global $Φ$ optimum rather than merely search for it: an exact moment-map zero fixes $Φ_{\min} = 200/9$, and de Groote's classification extends that optimality to every exact real rank-7 $2\times2$ decomposition.
Every exact real rank-7 realization therefore has a $Φ$-predicted RMS constant at least $5/3$ times that of the cubic algorithm, at fixed noise coefficient.
We then introduce an explicit block-scaled e4m3 model in which $Φ$ is the leading-order coefficient of relative expected mean-squared error, and we test the resulting $Φ$-predicted ordering against realized fp8 error.
Ordering and re-basing experiments support that prediction within tested fused block-scaled regimes, and on real matmul tiles from four architecture families the $Φ$-optimal realization falls in the fp8 low-error region.
Across two $\sim$70B models on real deep_gemm kernels, the same realization removes 10 to 55% of classic Strassen's excess NLL over the clean model.
Algorithm realization thus becomes a mathematically certified design problem rather than a tuning choice: an independent low-precision axis with a global $Φ$ optimum and measured fp8 relevance.