Heterogeneous multi-agent systems combine models with different capabilities through a common communication interface.
Exchanging internal states directly requires translating between model-specific representations and controlling intermediate computation.
We introduce the Vision Wormhole
Which repurposes the visual input interface of Vision-Language Models (VLMs) for continuous communication between frozen heterogeneous agents.
Universal Visual Codec
Encodes each sender's latent rollout into a fixed-size message, maps it through a shared reference space, and decodes received messages into the receiver's image-token span.
Per-model codecs and affine reference maps form a hub-and-spoke architecture with $O(N)$ components for $N$ models.
Each model learns its codec independently through self-distillation on anchor texts, and shared-anchor alignment enables reuse across communication partners.
Performance Improvements
Across four VLM families, six team configurations, and nine reasoning benchmarks, Vision Wormhole improves accuracy by 6.0 percentage points on average over text-mediated MAS and achieves a 1.69$\times$ geometric-mean speedup in batch-normalized end-to-end runtime.