Key points are not available for this paper at this time.
We developed a new computational algorithm for the accurate identification of ligand binding envelopes rather than surface binding sites. We performed a large scale classification of the identified envelopes according to their shape and physicochemical properties. The predicting algorithm, called PocketFinder, uses a transformation of the Lennard-Jones potential calculated from a three-dimensional protein structure and does not require any knowledge about a potential ligand molecule. We validated this algorithm using two systematically collected data sets of ligand binding pockets from complexed (bound) and uncomplexed (apo) structures from the Protein Data Bank, 5616 and 11,510, respectively. As many as 96.8% of experimental binding sites were predicted at better than 50% overlap level. Furthermore 95.0% of the asserted sites from the apo receptors were predicted at the same level. We demonstrate that conformational differences between the apo and bound pockets do not dramatically affect the prediction results. The algorithm can be used to predict ligand binding pockets of uncharacterized protein structures, suggest new allosteric pockets, evaluate feasibility of protein-protein interaction inhibition, and prioritize molecular targets. Finally the data base of the known and predicted binding pockets for the human proteome structures, the human pocketome, was collected and classified. The pocketome can be used for rapid evaluation of possible binding partners of a given chemical compound. We developed a new computational algorithm for the accurate identification of ligand binding envelopes rather than surface binding sites. We performed a large scale classification of the identified envelopes according to their shape and physicochemical properties. The predicting algorithm, called PocketFinder, uses a transformation of the Lennard-Jones potential calculated from a three-dimensional protein structure and does not require any knowledge about a potential ligand molecule. We validated this algorithm using two systematically collected data sets of ligand binding pockets from complexed (bound) and uncomplexed (apo) structures from the Protein Data Bank, 5616 and 11,510, respectively. As many as 96.8% of experimental binding sites were predicted at better than 50% overlap level. Furthermore 95.0% of the asserted sites from the apo receptors were predicted at the same level. We demonstrate that conformational differences between the apo and bound pockets do not dramatically affect the prediction results. The algorithm can be used to predict ligand binding pockets of uncharacterized protein structures, suggest new allosteric pockets, evaluate feasibility of protein-protein interaction inhibition, and prioritize molecular targets. Finally the data base of the known and predicted binding pockets for the human proteome structures, the human pocketome, was collected and classified. The pocketome can be used for rapid evaluation of possible binding partners of a given chemical compound. Prediction of ligand binding sites is a fundamental step in the investigation of the molecular recognition mechanism and function of a protein. An increasing number of protein structures are becoming available from high throughput structural genomic projects prior to biological and functional characterization. Therefore, computational methods to predict ligand binding sites are becoming increasingly important. There are three independent sources of information that can be used to infer the location of possible ligand binding sites on the surface of a protein: (i) protein structure, (ii) evolutionary information (sequence alignments), and (iii) ligand/substrate information. A number of sophisticated algorithms using evolutionary information or algorithms predicting locations of binding sites for specific substrates have been published (1Campbell S.J. Gold N.D. Jackson R.M. Westhead D.R. Ligand binding: functional site location, similarity and docking.Curr. Opin. Struct. Biol. 2003; 13: 389-395Google Scholar, 2Lichtarge O. Yao H. Kristensen D.M. Madabushi S. Mihalek I. Accurate and scalable identification of functional sites by evolutionary tracing.J. Struct. Funct. Genomics. 2003; 4: 159-166Google Scholar, 3Lichtarge O. Sowa M.E. Evolutionary predictions of binding surfaces and interactions.Curr. Opin. Struct. Biol. 2002; 12: 21-27Google Scholar). Here we attempted to develop an algorithm that is based solely on the protein structure and without any prior knowledge about the nature of the substrate. We hypothesized that the structure itself is sufficiently informative, whereas the evolutionary conservation and the nature of the ligand can only be used as optional contributions. Proteins are involved in several kinds of molecular interactions: with other proteins, DNA, RNA, peptides, and small molecules. In this study we present an algorithm to predict the binding envelopes near potential small ligand binding sites or areas that could be targeted with small “druglike” compounds. Once the ligand binding pocket is predicted, a high throughput ligand docking procedure or structure-based drug design (4Walters W.P. Stahl M.T. Murcko M.A. Virtual screening—an overview.Drug Discov. Today. 1998; 3: 160-178Google Scholar, 5Klebe G. Recent developments in structure-based drug design.J. Mol. Med. 2000; 78: 269-281Google Scholar, 6Gane P.J. Dean P.M. Recent advances in structure-based rational drug design.Curr. Opin. Struct. Biol. 2000; 10: 401-404Google Scholar, 7Abagyan R. Totrov M. High-throughput docking for lead generation.Curr. Opin. Chem. Biol. 2001; 5: 375-382Google Scholar, 8Shoichet B.K. McGovern S.L. Wei B. Irwin J.J. Lead discovery using molecular docking.Curr. Opin. Chem. Biol. 2002; 6: 439-446Google Scholar, 9Anderson S. Chiplin J. Structural genomics: shaping the future of drug design?.Drug Discov. Today. 2002; 7: 105-107Google Scholar) can be used to generate a list of the lead molecules. The properties of druglike molecules are well studied (10Veber D.F. Johnson S.R. Cheng H.Y. Smith B.R. Ward K.W. Kopple K.D. Molecular properties that influence the oral bioavailability of drug candidates.J. Med. Chem. 2002; 45: 2615-2623Google Scholar, 11Lipinski C.A. Drug-like properties and the causes of poor solubility and poor permeability.J. Pharmacol. Toxicol. Methods. 2000; 44: 235-249Google Scholar) and cover a certain range of sizes, typically with molecular mass between 300 and 700 daltons. Therefore, we excluded from consideration very small ligands, such as metals and small solvent molecules, as well as very large substrates. However we wanted to develop an algorithm that within this size range does not depend on the nature of the ligand. A number of structure-based pocket prediction algorithms have been published over the last 10 years. They can be divided into two general classes: (i) geometric algorithms and (ii) probe mapping/docking algorithms. Geometric approaches analyze protein surfaces to find clefts. SURFNET (12Laskowski R.A. SURFNET: a program for visualizing molecular surfaces, cavities, and intermolecular interactions.J. Mol. Graph. 1995; 13: 323-330Google Scholar) detects the gap regions in proteins by fitting spheres into the spaces between protein atoms. The sphere fitting process results in a number of separate groups of interpenetrating spheres, which correspond to the cavities and clefts of the protein. LIGSITE (13Hendlich M. Rippmann F. Barnickel G. LIGSITE: automatic and efficient detection of potential small molecule-binding sites in proteins.J. Mol. Graph. Model. 1997; 15: 359-363Google Scholar), an improved version of POCKET (14Levitt D. Banaszak L. POCKET: a computer graphics method for identifying and displaying protein cavities and their surrounding amino acids.J. Mol. Graph. 1992; 10: 229-234Google Scholar), identifies clefts by putting the protein in a regular Cartesian grid and scanning along the x, y, and z axes and the cubic diagonals for areas that are enclosed on both sides by protein. APROPOS (15Peters K.P. Fauck J. Frommel C. The automatic search for ligand binding sites in proteins of known three-dimensional structure using only geometric criteria.J. Mol. Biol. 1996; 256: 201-213Google Scholar) and CAST (16Liang J. Edelsbrunner H. Woodward C. Anatomy of protein pockets and cavities: Measurement of binding site geometry and implications for ligand design.Protein Sci. 1998; 7: 1884-1897Google Scholar) are based on the α-shape algorithm, which identifies pockets by comparing surfaces of the protein generated with different levels of detail. PASS (17Brady Jr., G.P. Stouten P.F. Fast prediction and visualization of protein binding pockets with PASS.J. Comput.-Aided Mol. Des. 2000; 14: 383-401Google Scholar) identifies the “active site points” by coating the protein surface with a layer of spherical probes and then filtering out those that clash with the protein or are not sufficiently buried. In addition to those pure geometrical methods, some approaches based on mapping/docking and scoring of molecular fragments have been proposed (18Dennis S. Kortvelyesi T. Vajda S. Computational mapping identifies the binding sites of organic solvents on proteins.Proc. Natl. Acad. Sci. U. S. A. 2002; 99: 4290-4295Google Scholar, 19Kortvelyesi T. Silberstein M. Dennis S. Vajda S. Improved mapping of protein binding sites.J. Comput.-Aided Mol. Des. 2003; 17: 173-186Google Scholar, 20Ruppert J. Welch W. Jain A. Automatic identification and representation of protein binding sites for molecular docking.Protein Sci. 1997; 6: 524-533Google Scholar, 21Verdonk M.L. Cole J.C. Watson P. Gillet V. Willett P. SuperStar: improved knowledge-based interaction fields for protein binding sites.J. Mol. Biol. 2001; 307: 841-859Google Scholar, 22Bliznyuk A. Gready J. Simple method for locating possible ligand binding sites on protein surfaces.J. Comput. Chem. 1999; 9: 983-988Google Scholar, 23Glick M. Robinson D.D. Grant G.H. Richards W.G. Identification of ligand binding sites on proteins using a multi-scale approach.J. Am. Chem. Soc. 2002; 124: 2337-2344Google Scholar). Two excellent recent reviews of computational tools for identification of small molecule binding sites in proteins give a good overview of the field (1Campbell S.J. Gold N.D. Jackson R.M. Westhead D.R. Ligand binding: functional site location, similarity and docking.Curr. Opin. Struct. Biol. 2003; 13: 389-395Google Scholar, 24Sotriffer C. Klebe G. Identification and mapping of small-molecule binding sites in proteins: computational tools for structure-based drug design.Farmaco. 2002; 57: 243-251Google Scholar). Pure geometric methods are relatively straightforward, but there is no direct physical meaning behind them. On the other hand, methods using molecular fragment mapping and ligand docking are better physically justified but computationally expensive and cannot always provide a good discrimination between correct and incorrect sites. Also the previously published methods were typically tested on relatively small data sets. Although APROPS used a relatively large test set of about 300 structures, others only used 10–50 selected test cases. This is despite the fact that almost 30,000 x-ray structures have been deposited in the Protein Data Bank (25Bernstein F.C. Koetzle T.F. Williams G.J. Meyer Jr., E.E. Brice M.D. Rodgers J.R. Kennard O. Shimanouchi T. Tasumi M. The Protein Data Bank: a computer-based archival file for macromolecular structures.J. Mol. Biol. 1977; 112: 535-542Google Scholar, 26Berman H.M. Westbrook J. Feng Z. Gilliland G. Bhat T.N. Weissig H. Shindyalov I.N. Bourne P.E. The Protein Data Bank.Nucleic Acids Res. 2000; 28: 235-242Google Scholar). Finally, because the ultimate goal of the binding site prediction methods is to find active sites on uncharacterized structures, it is important to test and validate the algorithms on large sets of the “unbound” or apo structures. Only the PASS algorithm was tested on a data set of 21 apo structures; the other publications did not test the effect of induced conformational changes on the prediction accuracy. A benchmark test based on a large, systematic data set of apo structures is necessary for evaluating protein-ligand binding site identification methods. In this study, we present and validate a novel algorithm for prediction of ligand binding pockets. The algorithm called PocketFinder is based on a transformation of the Lennard-Jones potential. Like pure geometric approaches, PocketFinder is fast and capable of identifying clefts and cavities regardless of the nature of the substrate while being more sensitive and specific. Furthermore, in contrast to other methods, the PocketFinder algorithm not only detects the location of the binding pocket but also predicts envelopes representing the shape and size of putative ligand binding volume. The method was tested on a systematically collected data set 2 orders of magnitude larger than previous benchmarks: 5,616 binding sites collected from ligand-protein complexes and 11,510 apo binding sites inferred from the complexes by homology. All small molecule binding envelopes from the human structural proteome were collected and clustered into a pocketome. The predicted binding envelopes were hierarchically clustered. The complete pocketome may be useful for understanding a complex network of interactions between small molecules and the cell proteins. All three-dimensional protein structures were taken from the October 3, 2003 Protein Data Bank release. Prior to computation, we removed all ligands and water molecules from the structure. Protein-ligand binding pockets are predicted based on the grid potential map of van der Waals interaction of the the regions of high van der Waals the procedure was The step was to the grid potential map of the van der Waals field using a probe for an were in surrounding M. R. protein-ligand docking by in 1997; Scholar). The grid and a of the of the protein was The potential was calculated according to the Lennard-Jones is the between the probe at a grid and the protein and are taken from the for molecular were into the range to only the The step was to the potential map to the regions with the van der Waals potential and to was performed by an of the potential on grid with the of this transformation a computationally with an were to the of The step was to putative ligand envelopes by the map at a calculated as the was on the data and are an and a of all map potential The step was to the envelopes by their and out those than All for the map were using a large, set of binding sites. The algorithm was with the used pocket set collected from pocket set collected from apo structures; overlap of predicted to the binding Structural of function R. Totrov M. D. a new method for structure and to docking and structure prediction from the Comput. Chem. 15: Scholar, Scholar). was to develop and validate an accurate algorithm for predicting protein-ligand binding sites to which a druglike small molecule may we studied the size of known We then a data base of protein ligand binding sites for the of the pocket prediction We from the October 2003 of the Protein Data Bank that structures of proteins. we collected all structures that are complexed with by in the Protein Data to a data set of binding sites from protein-ligand complexes set This in Protein Data Bank and were used to the protein-ligand binding site data we a ligand size and than were excluded in this This excluded metals and a benchmark the ligands, we excluded high or such as certain This the number of to of all a was Only with better than were Proteins that of than or more than were also This the number of to We other that did not the size of the data set but the data as (i) that are from the receptors of the were within from the were (ii) that the of the in were removed because their binding sites are between the and a correct biological (iii) were and of Protein Data Bank and ligands were The data set of 5,616 protein-ligand binding sites is the of Protein Data Bank with goal was to predict a potential binding from an uncomplexed structure, it was to a benchmark of pocket sites This data set to validate the pocket prediction algorithm in a more pockets may not be as as the pockets to the conformational a of the pocket in the of the ligand. The was collected by the protein structure in the uncomplexed and mapping the ligand to them. The proteins in the to the (i) The between apo structure and complexed protein is over (ii) The of the structure is better than (iii) There are no the surface within the ligand. There are no other ligands within the ligand in the apo structure. Only the receptors in the were used to find the apo structures. The only for using the receptors was to the process of identifying and the structures. We all binding pockets in the to an of A of 11,510 pockets were collected from the Protein Data Bank binding sites to binding sites from As a the on about different of the same binding site in We calculated version of the three-dimensional Lennard-Jones potential on a grid surrounding the protein surface and and then at an to pocket The pocket was by a surface We this potential because in contrast to geometrical methods it a physical and at the same it does not require any knowledge of the chemical nature of a ligand. the van der Waals of the binding is present in complexes of physical or predicting the location of potential binding we the of those The of prediction was by the overlap of protein in with the ligand and protein in with the predicted This overlap was calculated as is the of the within from a bound and is the of the within from the predicted A prediction have to whereas a prediction have to this be because the same pocket may different ligands, and all ligands may a site but in different We the of prediction as at 50% of the to be by the predicted We PocketFinder to the to evaluate on structures. The method identified 96.8% of the 5,616 ligand binding sites as a overlap between the and predicted sites the of the Furthermore of binding sites were identified and of binding sites were identified with than than overlap A more test was performed using a data set of 11,510 binding sites collected from uncomplexed structures We that to the of the binding pockets be 95.0% of the binding sites were predicted with than results suggest that the method is to the conformational of binding sites in uncomplexed structures. in the high only of the apo binding sites than as with in the the of binding sites it was with for the We studied the effect of different conformational in the binding sites of the the prediction the 11,510 binding sites of the were inferred based on binding sites in the we into and the by the of The of the of for is in this we the (i) The of (ii) the with high of the overlap of the with the but the were some larger of the (iii) the with of the was We better overlap for the than the some of pockets. for the of the pockets more to predict than the pockets. we may that the of receptors were not as good as the for the binding site there was only a of recognition the of the predicted envelopes between and apo structures, we algorithm to a set of with induced as by S. Identification of specific interactions that in with Mol. Biol. Scholar). we the of the proteins. with relatively small A and the envelopes predicted from both and apo are very and the bound the with large and the envelopes different and between the and apo structures. the ligand in was and the ligands in and were The algorithm may any number of putative binding envelopes on the nature of the protein surface and the of a predicted the between the protein size and the number of predicted pockets. The number of envelopes predicted protein is to the of the protein pocket for the of the proteins is or to in the data than envelopes were several envelopes are it is important to which to the binding in the In the of we that the envelopes with the ligand were the of the predicted envelopes with the ligand were the and were the We the envelopes by and the of This that the algorithm may several putative binding the two predicted binding sites cover as many as of the binding sites. and R.A. Protein clefts in molecular recognition and Sci. 1996; 5: Scholar) a based on a data set of The size of the predicted is important of the prediction because envelopes to binding site but are at to The PocketFinder was to generate envelopes that the binding sites. the of we calculated (i) the of the envelopes with to the of the bound ligands and (ii) the of the predicted binding to the surface of the for binding site in the The of the ligands was whereas the of the envelopes was On the predicted was larger in than the ligand. that ligands their pockets and the ligands bound to those receptors were some of all possible ligands, this the predicted binding the to the surface of the was the of the of to the surface of the protein for bound ligands and predicted This that the size of the predicted binding was to the binding with the high of the we can that the prediction was there are many pockets in the benchmark to prediction to a of the we an of the benchmark pockets using the We the proteins in the based on the Structural of Proteins data base A. D. C. in structure and Acids Res. Scholar). They to 10 of in high of data from and were in were selected from different based on the (i) the size between 300 and 700 the range of that was and (ii) the that only in were to any of proteins from We then by their and the only the and the it envelopes were the and of the is important to that in the predicted envelopes the ligands with size and shape being to the bound The complete data sets and and prediction results are available As the size of structural proteome we can a pocketome, as the of all possible small molecule envelopes present in a Here we and all envelopes from human protein-ligand complexes with known three-dimensional structure. There are binding pockets from human proteins in the the shape surface and the three and two physicochemical and J. A method for displaying the of a Mol. Biol. Scholar) and were used to the representing by a a was A algorithm was used to a from the We the of the human ligand envelopes with the based on the chemical similarity between This the complex between ligands and their pockets. We that the same pocket can different ligands of different size and and the same small molecule can to rather different pockets. a fragment of those from binding sites for all binding sites of available human protein-ligand complex structures is available is by the Protein Data Bank and ligand three different of the classification by the evolutionary the classification by the ligands chemical similarity of the bound ligands, and the classification by the envelopes the similarity of the binding pockets. We groups of ligands in the ligand by and and in of the the ligands to the same were clustered in the other as the between the was the complex nature of the a ligand can to different pockets, and the same pocket may be good for different The of putative ligand envelopes a new for understanding the the number of envelopes we can or a pocket is in the pocketome. We that the PocketFinder algorithm can or predict protein-ligand binding envelopes from complexed and apo structures. We also the of a human of this pocketome can along with the structural
An et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: