This review demonstrates state-of-the-art protein function prediction with pretrained transformers, highlighting model selection for biologists.
Transformer-based protein language models (PLMs) learn meaningful representations from millions of unlabeled sequences, capturing evolutionary patterns and functional relationships. Recent advances include ESM-2’s systematic scaling to 15 billion parameters, structure-aware vocabularies (SaProt), and multimodal foundation models (ESM-3, 98B parameters). PLMs achieve state-of-the-art performance: Gene Ontology prediction (F-max 0.64–0.68), enzyme classification (81% accuracy), and variant effect prediction (Spearman ρ 0.52–0.55). Deep-layer attention correlates 44–63% with 3D contacts despite no structural training. This review synthesizes recent PLM developments, benchmarks, and practical applications, providing guidance for experimental biologists on model selection and validation strategies.
No takes yet. Share an insight, caveat, or question.
Kushal Raj Roy (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: