Key points are not available for this paper at this time.
Abstract Cryptic binding sites in proteins, which are hidden in the absence of a ligand, offer opportunities to modulate targets previously considered ‘undruggable’. However, the scarcity of experimentally validated examples limits the development of predictive tools. Here, we introduce CryptoBank, a large-scale database of cryptic sites identified by applying a machine learning model to detect ligand-induced conformational changes in over 5.5 million structural alignments of unbound (apo) and bound (holo) protein pairs from the Protein Data Bank (PDB). Our analysis reveals that cryptic pockets are widespread, occurring in approximately 16.3% of protein clusters. Leveraging this resource, we fine-tuned a protein language model (PLM) to predict cryptic sites directly from protein sequence information. This sequence-based model achieves high precision (PR AUC 0.8) when query sequences share more than 20% identity with CryptoBank entries. Critically, we demonstrate its broader utility by predicting a cryptic site in human TPP1, a protein with less than 20% sequence identity to any CryptoBank entry, and validating its opening using molecular dynamics simulations. CryptoBank and the predictive PLM are publicly accessible via a web server, providing valuable resources for cryptic site discovery and drug development.
Martinez et al. (Wed,) studied this question.