首页 > AI前沿 > NeuralZip: Reusable Setup for Fast Lossless Compression

NeuralZip: Reusable Setup for Fast Lossless Compression

arXiv机器学习 2026-10-07 20:06 3 阅读 查看原文

Lossless compression can reduce the storage and movement of model weights without changing their floating-point values, but repeated statistical analysis and code construction add computational overhead.

We study whether the statistical structure of exponents can be prepared once and reused. For this, we introduce NeuralZip, which groups chunks with similar exponent distributions, shares Huffman codes, and selectively represents recurring exponent tuples using packed exponents, thereby achieving additional moderate compression ratios.

A setup chooses these representations before subsequent encodings, while every encoding still processes the current tensor values.

In floating-point model checkpoints, post-setup compression is 1.81-21.33× faster than the baselines and achieves exact bit-to-bit reconstruction.

We show that this setup can be precomputed and transferred from another compatible architecture, preserving similar compression ratios and avoiding the need to amortize setup costs.

Therefore, compression adaptation is transferable and reusable.

Training checkpoints demonstrate continued reuse as the weights evolve.

Finally, GPU experiments reduce active memory usage by up to 27.5% while reproducing the logits exactly.