首页 > AI前沿 > Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

arXiv机器学习 2026-10-06 12:51 5 阅读 查看原文

Adversarial training is one of the most reliable defenses against adversarial attacks, but its high computational cost must generally be paid anew for each task.

Robust foundation models offer a promising alternative: adversarially pretrain a model once and then transfer its robustness to downstream tasks through lightweight adaptation.

However, a fundamental question remains open: can robustness acquired during pretraining transfer to unseen tasks without further adversarial training?

In this study, we answer this question affirmatively.

A single model adversarially pretrained at scale can achieve optimal robustness on new tasks without additional task-specific training.

Specifically, we show that, for a family of Gaussian-mixture classification tasks, a sufficiently deep linear transformer adversarially trained across tasks can asymptotically attain the robust Bayes error on previously unseen tasks through in-context learning from clean demonstrations.

By contrast, a standardly trained model cannot.

We further analyze convergence under gradient flow, an accuracy--robustness trade-off, and demonstration complexity.