首页 > AI前沿 > Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

arXiv自然语言 2026-09-28 05:03 5 阅读 查看原文

A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error.

The estimator is well-behaved because round-to-nearest error is spectrally flat.

Experiments

Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit.

The probe coefficient of variation is predictable from tensor shape.

One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.

RAM Algorithm

RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data.

Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths.

A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments.

One probe pass serves any budget.

Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it.

That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective.

On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes.

Performance

Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models.

(Black Sheep Ai baa.ai)