Actor-critic methods achieve strong performance in continuous control, but their policies can produce highly oscillatory actions.
A common remedy is to add auxiliary smoothness losses.
However, their contribution can be negligible when their gradients are small relative to the native actor gradient.
Moreover, existing methods often combine multiple auxiliary losses, complicating loss balancing without necessarily improving the return-smoothness trade-off.
We introduce DAMPER (Direction-Aware Magnitude-Controlled Projection with Explicit Return Priority)
It combines the native actor gradient with a temporal-consistency gradient through conflict-conditioned projection and adaptive magnitude control.
It removes the auxiliary component opposing the actor gradient and scales the retained temporal direction relative to the actor gradient norm, preserving positive alignment with the native actor gradient.
Experiments
Experiments with TD3 and SAC on six continuous-control tasks show reduced action oscillation relative to the native agents in all 12 task-backbone pairs.
With task-dependent return trade-offs, the best oscillation score among the compared methods in eight.