首页 > AI前沿 > Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models

Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models

arXiv机器学习 2026-09-30 03:57 7 阅读 查看原文

Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives.

The premise that reducing these metrics improves the downstream understanding-generation trade-off has never been tested directly.

We audit it in a controlled testbed, GRIDUMM, which mirrors key structural ingredients of UMM training while making the ground-truth trade-off exactly computable.

Across 63 configurations and 372 measured checkpoints, no directional conflict metric reaches an absolute Spearman correlation of 0.3 with a confidence interval excluding zero for conflict measured during training against the eventual trade-off.

A dose-response intervention that monotonically suppresses conflict leaves the trade-off flat, separating correlation from causation.

The norm ratio is a generation-failure detector and becomes null among configurations that master generation.

Functional interference measures outperform directional conflict metrics, while training loss tracks the trade-off strongly.

Our results do not show that conflict is useless; they show that its validity as a diagnostic target must be established, not assumed, and we release the audit protocol as a reusable standard.