首页 > AI前沿 > PhyMo: A Physical-Field Modality for Multimodal AI4Physics

PhyMo: A Physical-Field Modality for Multimodal AI4Physics

arXiv机器学习 2026-09-23 16:41 6 阅读 查看原文

Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge.

However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions.

To address this limitation, we introduce the physical-field modality and propose PhyMo, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators.

PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads.

Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.