Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges.
We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity.
In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing.
Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error.
A post-hoc gauge fixing removes redundant parameters at no representational cost.
Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.