Autoregressive transformers remain comparatively weak for protein sequence and structure generation.
We study the role of target representation: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework.
Introducing Protein Latent Language (PLL) and Structure Latent Language (SLL)
We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences to a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, with one token per residue.
Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while retaining decoding to backbone coordinates.
Pretraining and Performance
We separately pretrain autoregressive transformer models on PLL and SLL tokens using next-token prediction, yielding PLLM and SLLM.
Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model.
In unconditional sequence generation, PLLM reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across sampling temperatures.
Sequence-to-Structure Prediction
For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training.
Performance on Long Proteins
For long proteins, latent-token sampling is approximately 1,000 times faster than MSA-based AlphaFold2 in our measurements.
Backbone Generation
In backbone generation, SLLM compares favorably with other generative models on diversity and novelty.
Inference-Time Sampling
We also observe early signs that using SLLM's internal token confidence for inference-time sampling can improve sequence-to-structure prediction quality beyond a single decoded sample.
Conclusion
These results position learned latent protein languages as a promising substrate for autoregressive transformer scaling and inference-time sampling in protein generation.