Protein language models, built on architectures originally developed for natural language processing (NLP), offer a powerful framework for representing protein sequences. We propose that such models can be enhanced by adopting strategies widely used in NLP. In particular, we explore in-context learning for protein language models (ICL4P) as a means to extend the utility of large pretrained models in low-data settings. Given the high cost of generating labeled biochemical data, methods that reduce data requirements are especially valuable for experimental biologists. As a proof-of-concept, we developed a peptide classifier based on ICL4P that can serve as a prescreening tool to prioritize peptide sequences before committing to costly high-throughput assays. Specifically, we show that ICL4P can be used to construct an effective classifier for predicting peptide binding to the MHC-II protein HLA-DR1 using minimal training examples. We demonstrate that this approach performs better than random chance with as few as three known peptide binder examples. Our results highlight the potential of in-context learning to make protein language models more accessible and practical in real-world biochemical applications.
Almonte et al. (Sun,) studied this question.