Models with near-identical predictive performance can yield substantially different predictions, a phenomenon known as predictive multiplicity.
Prior work has mostly studied this at the level of individual scalar outputs.
In time-series forecasting, however, predictions across horizons jointly define a trajectory, and horizon-wise comparisons can hide important differences in predictive behavior.
To address this problem, we introduce temporal predictive multiplicity, a framework that characterizes disagreement over complete forecast trajectories among models with near-identical predictive performance.
We show that constraining predictive performance alone can still admit a broad range of different trajectories.
We further show that constraining multiplicity at individual horizons partially reduces, but does not eliminate, trajectory-level multiplicity.
Experiments with 19 neural forecasting architectures on 11 datasets confirm that near-optimal models can exhibit substantial variability in the forecast trajectories they produce, and trajectory-level disagreement is largely unrelated to horizon-wise disagreement.
Our framework, therefore, exposes a gap in existing multiplicity studies:
models with indistinguishable predictive performance imply fundamentally different temporal trajectories, with consequential downstream effects.