首页 > AI前沿 > ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models

ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models

arXiv机器学习 2026-09-24 19:24 4 阅读 查看原文

Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector.

The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry.

Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity.

We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field.

Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods.

A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads.

We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score.

Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks.

ICE ranks first in all 30 reported supervised and few-shot comparisons.

Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.