Clinical diagnosis is inherently sequential: clinicians escalate from cheap to costly tests only when additional evidence is expected to resolve diagnostic uncertainty.
We present ActiveMedAgent, a framework that brings this cost-aware sequential logic to multimodal medical AI.
Given a frozen, API-accessed vision-language model, ActiveMedAgent tracks probability distributions over candidate diagnoses and scores each acquisition by its per-step diagnostic utility minus cost.
A lightweight MLP controller is then trained offline on these scored trajectories, learning when to request additional evidence and when to commit.
Experiments and Results
Across three commonly used benchmarks, trajectory-based policy learning consistently outperforms both unguided acquisition and full-modality baselines.
Notably, we identify an information overload effect.
In 175 cases, the agent produces a correct diagnosis with fewer channels while the full-modality baseline fails, showing that learning what to omit can be as important as learning what to acquire.