首页 > AI前沿 > TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference

TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference

arXiv自然语言 2026-08-01 18:38 5 阅读 查看原文

Mixture-of-experts (MoE) models improve capacity with moderate overhead by sparsely activating experts per token.

However, deploying MoE across resource-constrained edge servers incurs substantial cross-server communication as experts are distributed across heterogeneous servers.

Existing placement methods optimize for raw token traffic, while conventional compression considers semantics but ignores topology-dependent routing costs.

Consequently, independent optimization leads to inefficient communication and resource utilization.

This paper proposes TopoCompress

a deployment- and topology-aware token compression framework for communication-efficient distributed edge MoE inference.

It jointly optimizes token compression, expert deployment/replication, GPU-CPU residency, and collaborative routing to balance cross-server transmission, quality, and resource use.

To address the coupling between token-level compression and epoch-level deployment, TopoCompress employs a two-timescale alternating optimization.

In the online fast loop, it identifies and compresses low-importance, high-routing-cost tokens and jointly routes surviving expert activations.

In the offline slow loop, it updates expert placement, replication, and GPU-CPU residency according to post-compression traffic accumulated during online inference.

We establish the feasibility, optimality, convergence, and computational complexity.

Simulations demonstrate that TopoCompress effectively reduces cross-server traffic and deployment resource consumption while maintaining controllable inference quality, enabling efficient distributed MoE inference over bandwidth- and resource-constrained edge infrastructures.