Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback.
Controlling such systems from high-dimensional visual observations requires representations that preserve these modes.
Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data.
However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible.
To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control.
We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace.
Thus, the encoder cannot discard any state direction reachable by an action sequence within $H$ steps.
Moreover, we show that, as $H$ grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace.
We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.