首页 > AI前沿 > Architecture-Dependent Fusion Pathways in MLLMs

Architecture-Dependent Fusion Pathways in MLLMs

arXiv机器学习 2026-10-02 21:32 6 阅读 查看原文

Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood.

We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures.

We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces.

Separately, we perform causal intervention experiments as a validation of the resulting interpretation.

As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis.

Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization.

This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.