首页 > AI前沿 > Learning Latent Protein Languages for Autoregressive Generation

Learning Latent Protein Languages for Autoregressive Generation

arXiv机器学习 2026-10-03 03:44 6 阅读 查看原文

Autoregressive transformers remain comparatively weak for protein sequence and structure generation.

We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework.

Introducing Protein Latent Language (PLL) and Structure Latent Language (SLL)

We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue.

Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates.

Pretraining and Performance

We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM.

Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model.

In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures.

Sequence-to-Structure Prediction

For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training.

Performance on Long Proteins

For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements.

Backbone Generation

In backbone generation, SLLM compares favorably with other generative models on diversity and novelty.

Inference-Time Sampling

We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample.

Conclusion

These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.