首页 > AI前沿 > Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion

Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion

arXiv机器学习 2026-10-06 06:25 5 阅读 查看原文

Multimodal systems often encode every available input, even when a subset suffices for prediction.

Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop.

However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set.

We introduce SemARC

SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches.

We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders.

ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning

to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order.

Performance Results

  • Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset's most accurate baseline.
  • End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average.
  • Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average.

Thus, SemARC offers a practical path toward efficient multimodal inference across heterogeneous devices.