Large language models remain vulnerable to jailbreaks, and automated red teaming is the standard way to find jailbreaks in large language models at scale.
Current methods either draw more samples at test time through search, rewriting, and tree expansion, or train a stronger attacker offline with reinforcement learning.
Both share a limitation: once an attack on a specific target behavior begins, the attacker's weights are frozen. Any signal it gathers about the behavior stays in its context window and is discarded afterward. The attacker never adapts its proposal distribution mid-attack, so success depends almost entirely on the sampling budget, and under a budget affordable at scale, many behaviors remain unbroken.
We propose Red-TTT, which updates the attacker's parameters during the attack on each behavior.
At each round, the attacker samples a group of candidates, scores them against the victim's replies, and takes a policy-gradient step before drawing the next group, so what it discovers about the current victim is consolidated into weights rather than accumulated as context.
We also adapt the training objective to red teaming, where success is judged by the single best sample rather than the average.
Red-TTT requires only sampling access to the victim and integrates into existing attack pipelines with no other changes.
Against the Best-of-N baseline, Red-TTT raises attack success rate from 55.9% to 72.4% on average at a budget of 120 samples, improving over the baseline in every configuration and cracking many behaviors previous method cannot.
The code is available at https://github.com/SaFo-Lab/Red-TTT.