Intrinsically disordered proteins and regions (collectively IDRs) are found across all kingdoms of life and play critical roles in virtually every eukaryotic cellular process. In contrast to folded proteins, IDRs lack a stable 3D structure and are instead described in terms of a conformational ensemble, a collection of energetically accessible interconverting structures. This unique structural plasticity facilitates diverse molecular recognition and function; thus, a convenient way to view IDRs is through their ensembles. Here, we introduce STARLING, a latent denoising-diffusion model that combines a variational autoencoder (VAE), a transformer-based protein sequence encoder, and a Vision Transformer (ViT) to directly generate IDR ensembles from sequence. Leveraging modern multi-scale generative modeling, STARLING can generate ensembles of hundreds of conformers in seconds on both GPUs and CPUs. By jointly training the sequence encoder and the diffusion model, STARLING learns sequence features that determine a sequence’s ensemble in the form of protein embeddings. These embeddings enable the ensemble-first design and optimization of sequences within seconds, a task that traditionally requires extensive computational resources and considerable time. Furthermore, the embeddings support fast, large-scale searches for disordered protein ensembles directly from sequence, allowing users to query by conformational behavior rather than sequence identity alone. This unique combination of rapid ensemble generation and ensemble-aware protein embeddings dramatically lowers the barrier to the computational interrogation of IDR function through the lens of emergent biophysical properties, complementing bioinformatic protein sequence analysis. We evaluate STARLING’s accuracy against extant experimental data and offer a series of vignettes illustrating how STARLING can enable rapid hypothesis generation for IDR function and aid the interpretation of experimental data.
Novak et al. (Sun,) studied this question.