During the past two decades, mass spectrometry has become established as the primary method for protein identification from complex mixtures of biological origin. This is largely attributable to the fortunate coincidence of instrumental advances that allow routine analysis of minute amounts (typically femtomoles) of involatile, polar compounds such as peptides in complex mixtures, with the rapid growth in genomic databases that are amenable to searching with mass spectrometry (MS) 1The abbreviations used are: MS, mass spectrometry; MALDI, matrix-assisted laser desorption/ionization; ESI, electropray ionization; TOF, time-of-flight; LC, liquid chromatography; MS/MS, tandem MS; QqTOF, quadrupole mass selector and quadrupole collision cell with orthogonal acceleration TOF; CID, collision-induced dissociation; HPLC, high-pressure LC; ICAT, isotope-coded affinity tag. data. Like many other developing fields in science, the creation of techniques and software tools and the initial generation and interpretation of data have been the domain of experts, people who are cognizant not only of the benefits of the methods but also of their actual and potential weaknesses. Now, as mass spectrometric techniques and proteomic tools become increasingly available and accessible, a much broader range of researchers is applying the same methodology, often with substantially less understanding of the major limitations that critically affect the reliability and significance of the results. Ideally, the MS community should establish criteria for mass spectrometric identification of proteins that should be employed by researchers. As this remains a rapidly developing field with many different experimental approaches and different ways of searching and interpreting the data, it is difficult to promulgate hard and fast rules. Nevertheless, Molecular & Cellular Proteomics is attempting to develop standards of acceptability for proteomics papers, based on emerging knowledge as well as on principles of biological MS established over the last 20 years by the MS community. Authors of proteomics papers employing MS must make themselves fully aware of the key issues that are driving development of these guidelines. Hence, the paper that follows attempts to highlight the strengths and weaknesses of the methods in current use. It is particularly important to realize that for any protein match returned from a database search, there is a non-zero probability that it will be wrong. Many times, the quality of the data is such that the probability of a false positive can be disregarded, but in some cases the identifications returned by the search engines are very likely incorrect. Therefore, it is unacceptable to simply list all the hits that come back from any search engine and then discuss their biological significance as though they were categorically correct. Almost without exception, protein identification is based on the analysis of peptides generated by proteolyic digestion. The most widely used enzyme is trypsin, which hydrolyzes the protein on the C-terminal side of lysine and arginine, unless the subsequent amino acid in the sequence is a proline. This is advantageous as every peptide other than the protein C terminus has at least two sites for efficient protonation, the N-terminal amino group and the C-terminal basic residue, so peptides are readily ionized and detected as positive ions. However, for a variety of reasons, it is normal for only a subset of the potential tryptic peptides from any protein to be detected, particularly when the peptides are ionized directly from unseparated mixtures in which there may be competition for the available protons. There are also experimental limitations for the detection of peptides that are either very small or very large, a factor that is not controllable with a single enzyme as this is dependent on the distribution of lysine and arginine residues within a protein. Protein sequence coverage can be improved by carrying out a parallel digestion with a second protease of different specificity such as chymotrypsin, then combining the two digests for a common analysis. 2S. A. Carr, personal communication. In practice, whether a large or small fraction of the peptides generated from any protein is detected depends on many variables: the amount of that protein present in the original sample; the efficiency of any protein extraction, digestion, and peptide extraction; the presence of other proteins and other impurities; and the sensitivity and performance characteristics of the mass spectrometer and its mode of ionization, mass separation, and ion detection. Mass spectrometers employed in proteomic analysis use either matrix-assisted laser desorption/ionization (MALDI) or electrospray ionization (ESI) but they vary widely in their operation and performance characteristics. Early database searching was based on low-resolution linear MALDI-time-of-flight (MALDI-TOF) giving a mass accuracy of perhaps ±2 Da (1Henzel W.J. Billeci T.M. Stults J.T. Wong S.C. Grimley C. Watanabe C. Identifying proteins from two-dimensional gels by molecular mass searching of peptide fragments in protein sequence databases..Proc. Natl. Acad. Sci. U. S. A. 1993; 90: 5011-5015Google Scholar). However, this is no longer acceptable because good mass accuracy increases the reliability of database searching by limiting the possible compositions of peptides for any given mass (2Clauser K.R. Baker P. Burlingame A.L. Role of accurate mass measurement (± 10 ppm) in protein identification strategies employing MS or MS/MS and database searching..Anal. Chem. 1999; 71: 2871-2882Google Scholar). A delayed extraction MALDI-TOF instrument with reflectron should give better than 50 ppm mass accuracy, and significantly better (∼10 ppm) may be achieved with careful internal calibration. ESI has been the standard ionization method for liquid chromatography (LC)-MS and LC-tandem MS (MS/MS), although separated fractions can be deposited onto a MALDI target for either on-line or off-line analysis (3Preisler J. Hu P. Rejtar T. Karger B.L. Capillary electrophoresis-matrix-assisted laser desorption/ionization time-of-flight mass spectrometry using a vacuum deposition interface..Anal. Chem. 2000; 72: 4785-4795Google Scholar, 4Preisler J. Hu P. Rejtar T. Moskovets E. Karger B.L. Capillary array electrophoresis-MALDI mass spectrometry using a vacuum deposition interface..Anal. Chem. 2002; Scholar, identification of proteins with a ion mass Chem. Scholar). ESI is employed for single and and quadrupole ion that give The of a quadrupole mass selector and quadrupole collision cell with orthogonal acceleration and perhaps ppm mass accuracy well acceleration of mass Mass Scholar, ionization of mass of the Mass Scholar, time-of-flight mass spectrometer with electrospray ion and orthogonal Chem. Scholar, T. A. J. J. sensitivity tandem mass spectrometry on a time-of-flight mass Mass Scholar). MS the performance with mass accuracy of perhaps but such are and their has been to large MS There are at least approaches to protein identification based on peptide analysis. The is to as peptide mass or mass (1Henzel W.J. Billeci T.M. Stults J.T. Wong S.C. Grimley C. Watanabe C. Identifying proteins from two-dimensional gels by molecular mass searching of peptide fragments in protein sequence databases..Proc. Natl. Acad. Sci. U. S. A. 1993; 90: 5011-5015Google Scholar, A. T. T. T. of a A by by molecular ion mass Scholar, of protein Scholar, P. E. Protein identification by mass 1993; Scholar, P. P. of mass spectrometric molecular to proteins in sequence Mass 1993; Scholar, P. identification of proteins by 1993; Scholar, S. T. mass A to protein 1993; Scholar). This a of the MS mass with the molecular mass of the peptides generated by a digestion of protein in a Early with small databases as as or peptide be to the with mass from linear MALDI-TOF (1Henzel W.J. Billeci T.M. Stults J.T. Wong S.C. Grimley C. Watanabe C. Identifying proteins from two-dimensional gels by molecular mass searching of peptide fragments in protein sequence databases..Proc. Natl. Acad. Sci. U. S. A. 1993; 90: 5011-5015Google Scholar). genomic databases have as of the database which was than the criteria for protein identification have become and accurate mass measurement is It is also to match a of peptides and to a of the protein The second collision-induced of peptides from to the development of database peptide MS/MS from performance with fast ionization were used for peptide by tandem mass spectrometry of in Scholar). In the such MS/MS generated by MALDI or ESI, were sequence for all proteins in a of that be to of amino to of the peptide identification of peptides in sequence databases by peptide sequence Chem. Scholar, A.L. to tandem mass data of peptides with amino acid in a protein Mass Scholar, of the ion Chem. Scholar, K.R. Baker Burlingame A.L. from for searching of genomic Mass as in such as Protein In their such are based on the and all fragments for all peptides of the molecular based on rules. peptide match can be to a protein and it is possible that a single peptide will a protein although be in peptide the the of peptides to any protein and the the sequence the the probability of a searching will allow a peptide to be when it from a database peptide by perhaps a single amino and techniques have been to sequence although these are and database peptide by tandem mass Mass Scholar, Burlingame A.L. of the from T. using mass spectrometry and Chem. Scholar). no are because the protein is not present in the peptide to be a based on for peptide of peptides by tandem mass spectrometry and collision-induced Scholar). This is with good quality and is when a of peptides from to in genomic sequence and or and enzyme of the benefits of was the of the primary of J. Hu P. Rejtar T. Moskovets E. Karger B.L. Capillary array electrophoresis-MALDI mass spectrometry using a vacuum deposition interface..Anal. Chem. 2002; S. Burlingame A.L. of by mass spectrometry sequence analysis and molecular for a protein in the Chem. Scholar). The use of MS/MS and is the standard for protein identification and is peptide mass although the quality of tandem data with instrument ion from MALDI-TOF using E. by matrix-assisted mass Mass is to and the accuracy of mass is it is used particularly as MALDI has become available on performance A. peptide by a of and a mass Mass Scholar, of matrix-assisted laser desorption/ionization a time-of-flight spectrometer a Mass Scholar, matrix-assisted laser desorption/ionization and electrospray mass spectrometry for protein Mass 2000; Scholar, A. A. A. MALDI quadrupole time-of-flight mass a for proteomic Chem. 2000; 72: Scholar, Burlingame A.L. laser desorption/ionization with acceleration time-of-flight mass spectrometry for protein identification and Chem. and P. A delayed extraction MALDI-TOF for of protein Mass in and Scholar, P. Burlingame A.L. The characteristics of peptide collision-induced using a performance tandem mass Chem. 2000; 72: Scholar, A. P. J. A. A MALDI mass spectrometer for Chem. ion have been very and are to but different when in This is a as are separated not by mass but by and ESI ions. searching based on such low-resolution ion data based on the that the can have or some search all fragments are to be the probability that any peptide match will be incorrect. However, the of ion has a for software development and search engines that from also be ion may also be from analysis of ion J. J. T. A. peptide Scholar). are readily from the with and MS has for molecular and of proteins of to E. of proteins by mass Chem. 2002; Scholar). It is also important to that ion linear much in with other mass will the of a generation of tandem A largely factor in database searching is the of search of which there are in common use J. for MS protein Chem. 2002; Scholar). A of sites to database searching for peptide mass and the identification of sequence all of which other tools as such as for enzyme digestion and of sites allow of the to is to a to have the and some mass spectrometer software and software with the of of all the current search engines is a but tools on the proteomics by the of from and Protein by the of by instrument from and from Protein identification by digestion and peptide mass is not for complex protein mixtures unless by a most often a two-dimensional large it is possible to and or of protein from a single This is and but methods have been C. E. S. P. S. A. A. A molecular to proteomic and to Chem. 1999; 71: and for digestion, and of the peptide mixtures onto MALDI In practice, is to MS/MS, but is to a complex and then to the peptides to the mass mixtures may of proteins and high-pressure may be initial by chromatography may be by chromatography analysis of the by protein identification Scholar). mass is no longer possible with this as the is peptides and the proteins from which they were and peptide is The method is than two-dimensional in many it is the use of It can also be with in which that may proteins are and the peptides are to the mass spectrometer Baker J. P. Burlingame A.L. The identification of of the complex of S. using tandem mass 2002; Scholar). In with to a tandem mass of MS is with the and analysis of for analysis A. of liquid to the of in A. Scholar). The amount of data generated by such on a is for is of mass spectrometric data to the is dependent on the performance and accuracy of so are often to the instrument although some approaches have been for protein 2000; Scholar). Ideally, they the and mass of mass and some to the current and The for is by its of mass of the a fraction of the such as a that is to the of at the and to allow some to fully are but ESI are separated by A is to the in peptide of molecular mass less than Da with good this is as the in the Burlingame A.L. The and of the mass and mass accuracy for measurement of Mass in and is the most Da this is no longer and the identification of the difficult with particularly it a to actual and This is by the likely to in the of peptide the experimental data have been to possible the and to criteria that the data must then be as by the search engines allow for the of a of such as protein molecular mass mass for in a normal mass for mass for of of peptides for a and possible to residues such as of or of The then a of to of a of possible there are no standards for the from these is a match is not The by the search engines is the within Protein has been based on which for the distribution of peptide that from P. identification of proteins by 1993; Scholar). This is not particularly but some researchers have a of at least for a from peptide mass to be that no can be which a match is and which it is the of are developing a for Protein based on the of ion in MS/MS such as and than such as and that A.L. to tandem mass data of peptides with amino acid in a protein Mass protein identification by searching sequence databases using mass spectrometry 1999; and a that mass protein and data in a 2002; are to significance but to the may well be incorrect. A of have these by developing tools that the for search particularly of protein identifications using a Chem. 2002; Scholar, A. E. to the accuracy of peptide identifications by MS/MS and database Chem. 2002; Scholar, A for the of peptide in of peptide MS/MS and have tools A to peptides by of in Chem. Scholar, J. A. T. J. tandem mass spectrometry data Scholar, for tandem mass Chem. Scholar, A. E. A for proteins by tandem mass Chem. Scholar). Ideally, the will be of the mass spectrometer used and the search engine approaches on improved peptide and they a to a knowledge of peptide rules. of most of these methods is to make and to the for of data and its should the analysis of from with perhaps of of MS/MS method to a probability to peptide analysis of the peptides from a single protein. A of peptides for a given protein increases the probability of that protein based on protein identifications based on single peptides less than based on peptides A. E. A for proteins by tandem mass Chem. Scholar). has also been given to the in which a peptide sequence is of than as in many enzyme that for any analysis of a from a single there will be mass spectrometric that are not by the search of these may be as by of mass that for the compositions of normal peptide ions. may be peptides from the protein but from or carrying or However, in very many a from a single or a fraction to a single in will and only a subset of peptides will match any given protein peptide mass is the peptides that match to the can be then the list can be for second and subsequent However, MS/MS is better to peptide and and can also proteins will be detected when MS are with separation, either on-line or in because this a range and because it the ionization of only a subset of the peptides present in a In the reliability of protein identifications by MS, a variety of of must be there is by other proteins the initial and digestion of the which is difficult to for from the from or the as from there is the efficient extraction of the peptides and of the for MS to the presence of or that may affect The analysis should be for the a complex should not be without a either to analysis or on-line as in the are the mass method is the instrument is well and are the data should be as good as can In the data to a search the must a database and must that all the are possible and particularly the is not very with the search should be using different for the same and using different search the of hits from and Protein protein hits with search engines and may become because of peptide hits based on single peptides may also the other search engine a different which can as some different ion with Protein or the same In are also less likely to hits that are by only search it is not to every but a of should be particularly any that particularly important or In the should be the presence of should be and the should be to establish that and mass are correct. However, this will only be the data can be than or of the The from search will on the of the the techniques and the search but there are some principles to peptide mass there should be sequence coverage for the peptides the of peptides of the sequence not a and should not be as a positive In peptides of the sequence may be for a this there is no so should on the side of and in the of databases and the of or it should be to data from at least protein identifications based on data, it is important to that the peptide this is achieved can protein identifications have any Ideally, of a peptide match should that the to be the within the the range than all from the ions. there will be some to the and the for and in ESI, C-terminal arginine than ion N-terminal to a is it must be that the protein is a that may of a of it is this will be only the terminus is in a and then it is important to that this peptide to a will only be established amino acid sequence to the actual is in a or sequence can be based on the to a such a peptide may from but simply not be detected, it may have mass to or or the residues may be present in the sequence but the may not as It is important to that the of for the presence of a protein not establish the of the protein. though a protein may be in but not it is not to this as present in A but not This is that simply be by there is the of most mass spectrometric were small of In such it be to any identification by and a the creation of a be to whether there was any biological which for the identification of cell as a of the protein not to be the of cell to the Scholar). are in proteomic researchers have a to either or to present the of the search so that can make their many in proteomics only that proteins be can also be particularly for such as in of from and to the of MS, of the of of on two-dimensional gels by was a common to this The use of gels with mass spectrometric analysis and was by some C. E. S. P. S. A. A. A molecular to proteomic and to Chem. 1999; 71: Scholar). Nevertheless, this to may from range for in protein and the that the of a be to the presence of a protein than of the protein to be it is to use mass spectrometric analysis of that with the identification of proteins in mixtures may A single method for in protein which is available as the the use of may and to on a single two-dimensional This software that and protein to A different on the of MS to and researchers have proteins from in different or of the peptides from that be from the or by their mass of protein and Natl. Acad. Sci. U. S. A. 1999; Scholar, A. by amino in cell as a and accurate to 2002; Scholar). used in the MS analysis of peptides for at least 20 years is of peptide internal standards for field mass spectrometry of peptides in Mass Scholar). of a protein in in the at the C terminus of peptide with the of the protein C The use of a of and in mass for all fragments a peptide C to be from and other N-terminal proteomics using digestion in was given a with a method as carrying out two parallel but then from the only peptides from proteins in to be mass spectrometry for the rapid identification of Chem. 2002; Scholar). A factor is of by the protease that can to the presence of two than However, this can be from C. C. of of peptide and is advantageous as it a mass of Da than of the method to is to a amino acid such as with a or may isotope-coded affinity to of the as well as to measurement of analysis of complex protein mixtures using isotope-coded affinity 1999; Scholar). The of the is that it to be two of proteins that can then be any extraction, digestion, peptide separation, and analysis is As as peptides the different have and the of or from MS will be to the amounts of protein. In some are better than in and affect the in the of in peptide has no J. Burlingame A.L. Mass spectrometric analysis of protein mixtures at using affinity and Scholar). there is a distribution of in all it is to the in to the for the peptides with and The original used and a software for is available from instrument and on the that MS proteomics In the of for peptide can be to within although peptides or by may well is to and can be with methods that protein identification E. S. A. J. The of software tools to protein isotope-coded affinity and tandem mass of tandem mass spectrometry for protein and the of tools for data analysis and as has been for a from proteins E. S. A. J. The of software tools to protein isotope-coded affinity and tandem mass for peptide and proteins the of and tandem mass spectrometry to proteins with cell Scholar). that the methods only protein a such as in which amounts of peptides are to the spectrometric of with Chem. Scholar). peptides are to the peptides by of the proteins of and can to proteins such as J. of proteins and from cell by tandem mass Natl. Acad. Sci. U. S. A. Scholar). MS is for the of as all such a in molecular which is in the mass of any peptide carrying the amino this is attributable to the such as Da for or Da for are such as Da for the of but it may be whether this is the common of or It is also whether this is a or a mass may be Da or although these may be by the detection of at and a very accurate mass MS/MS techniques have been for the identification of the for peptides that such as or detection of in complex mixtures by electrospray liquid Mass 1993; Scholar, detection and of at the by mass Scholar, of by electrospray ionization and for detection of in protein Chem. 1993; Scholar). may be as with although of the group can give a peptide and to acid The major with the search for of such as in a methods have been employed to such as ion chromatography but with for of in using a affinity and matrix-assisted laser mass Sci. Scholar, P. affinity chromatography of Chem. 1999; 71: Scholar). The are when the is either or in the mass and MS/MS of the peptide should the of but a of acid from the molecular ion may the detection of any ions. In peptide mass be used for identification of on MS data a search can be for mass that are Da in mass than for the amino acid sequence but MS/MS to the sequence is the of be techniques for detection are Burlingame A.L. of in and of and by laser desorption/ionization (MALDI) and MALDI mass Mass 1999; and for of and proteins with a Scholar, Burlingame A.L. of sites of of factor using quadrupole time-of-flight mass Scholar). It is that the field of proteomic analysis by MS to be in a very it difficult to standards for and subsequent Nevertheless, to Molecular & Cellular Proteomics must that acceptable protein identification based on peptide mass should no longer be acceptable and must be by and it is that researchers the and for understanding the methods that they are it is that the of much in their their data, and its It will then be on and to fraction of this is and is for Authors should the software they the the database and the probability to protein Molecular & Cellular Proteomics will also the of that to experimental techniques or approaches to database searching and analysis that the accuracy and reliability of protein analysis based on mass spectrometric the methods in use can very to their has been largely The of this it that researchers in the field and to a and to protein
No takes yet. Share an insight, caveat, or question.
Michael A. Baldwin (2003) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: