Mass measurement is the main outcome of mass spectrometry-based proteomics yet the potential of recent advances in accurate mass measurements remains largely unexploited. There is not even a clear definition of mass accuracy in the proteomics literature, and we identify at least three uses of this term: anecdotal mass accuracy, statistical mass accuracy, and the maximum mass deviation (MMD) allowed in a database search. We suggest using the second of these terms as the generic one. To make the best use of the mass precision offered by modern instruments we propose a series of simple steps involving recalibration of the data on “internal standards” contained in every proteomics data set. Each data set should be accompanied by a plot of mass errors from which the appropriate MMD can be chosen. More advanced uses of high mass accuracy include an MMD that depends on the signal abundance of each peptide. Adapting search engines to high mass accuracy in the MS/MS data is also a high priority. Proper use of high mass accuracy data can make MS-based proteomics one of the most “digital” and accurate post-genomics disciplines. Mass measurement is the main outcome of mass spectrometry-based proteomics yet the potential of recent advances in accurate mass measurements remains largely unexploited. There is not even a clear definition of mass accuracy in the proteomics literature, and we identify at least three uses of this term: anecdotal mass accuracy, statistical mass accuracy, and the maximum mass deviation (MMD) allowed in a database search. We suggest using the second of these terms as the generic one. To make the best use of the mass precision offered by modern instruments we propose a series of simple steps involving recalibration of the data on “internal standards” contained in every proteomics data set. Each data set should be accompanied by a plot of mass errors from which the appropriate MMD can be chosen. More advanced uses of high mass accuracy include an MMD that depends on the signal abundance of each peptide. Adapting search engines to high mass accuracy in the MS/MS data is also a high priority. Proper use of high mass accuracy data can make MS-based proteomics one of the most “digital” and accurate post-genomics disciplines. Most analytical chemistry textbooks warn that data of unknown accuracy are useless. Yet many proteomics researchers entering the field in the last few years were raised on the idea that mass accuracy is an unimportant parameter. Until recently, much proteomics research was done on three-dimensional ion traps, very sensitive and robust but low resolution instruments. Database searches were performed with windows of ±3 Da and for several alternative charge states. Modern instruments can do more than a thousand times better and are capable of mass accuracy around 1 ppm. Speaking of masses here and throughout the text, we mean of course m/z values. Contrary to the old-fashioned magnetic sector instruments that imposed a compromise between mass accuracy and sensitivity this high accuracy now comes “for free.” That is to say modern instruments can be both extremely sensitive and capable of very high mass accuracy. In fact, the high resolution achieved by these instruments concentrates the signal into a narrow mass range, improving signal to noise of the spectra. Here we argue that proteomics thinking has not caught up with these capabilities and that consequently we are not making the best use of high mass accuracy. Surprisingly although the mass is the primary parameter measured in the mass spectrometric experiment, the proteomics community has not agreed on clear definitions. The proteomics literature attaches at least three different meanings to the term mass accuracy. This refers to the selective reporting of mass measurements, usually to demonstrate the capabilities of the author’s instrument. The literature is full of claims of very accurate measurements made on intrinsically not-so-accurate instruments, backed up with a single figure. Even for the highest resolution instruments, of the FTICR type, premature claims had been made of low ppm or sub-ppm mass accuracy. In practice, such performance in high throughput applications had to await solution of the space charge problem in FTICR, which only happened a few years ago. It is legitimate to report anecdotal mass accuracy, but it should clearly be distinguished from routine instrument performance in day to day use. In general, when reporting a single measurement we suggest to use the term “mass deviation,” defined as measured mass minus calculated mass, instead of the term mass accuracy. In other words, one should say “the mass was measured with a deviation of x ppm” instead of “the mass was measured with an accuracy of x ppm.” This is the mass accuracy estimated from a statistical distribution of mass errors. For example, the root mean square deviation of a particular mass may be reported by the manufacturer of a mass spectrometer as the mass accuracy of the instrument. The average absolute mass accuracy is a parameter easily calculated for any proteomics experiment and one that we have found to be quite informative of the quality of a measurement. A problem with the statistical mass accuracy may occur if the mass error does not follow the expected normal distribution. For instance, if the error distribution is Lorentzian instead of Gaussian the actual data dispersion σ is infinite, and thus any single measurement is potentially quite inaccurate. Moreover in this case statistical error cannot be reduced by averaging n independent measurements because σ/(n − 1)1/2 is still infinite. It is thus important to make sure (e.g. by the χ2 test) that the assumed Gaussian shape of the error distribution is valid. Plotting a graph of the error distribution and fitting to it a Gaussian distribution can be useful. Such a graph together with the absolute average deviation and dispersion σ are excellent indicators of actually achieved mass accuracy (see Fig. 1 for an example). This is the cutoff value used in database search. Only peptide sequences with a calculated mass within this tolerance are reported as hits. Maximum mass deviation (MMD) 1The abbreviation used is: MMD, maximum mass deviation. is the only operationally important parameter related to mass accuracy in a proteomics experiment. Inevitably in the proteomics literature anecdotal mass accuracy is the highest followed by the statistical mass accuracy, whereas MMD is usually chosen at several times the statistical mass accuracy or on a general feeling what the worst mass deviation is likely to be. We propose to use the generic term mass accuracy for the second of these three categories. Anecdotal mass accuracy is subjective, and MMD should be chosen as a multiple of the dispersion of mass errors; thus statistical mass accuracy is the only objective and independent parameter. As there are different ways to model the measured mass distribution, the statistical model should also be mentioned, for example root mean square error or absolute average deviation. If mass errors were normally distributed, three standard deviations would capture 99.7% of the peptides. For example, a root mean square error of 2.5 ppm would lead to an MMD of 7.5 ppm. However, for low signal to noise peaks the distribution is broader, and mass accuracy is worse than the average value would suggest. For these reasons and because the systematic error is not eliminated, many researchers set the MMD very high, in effect degrading a high accuracy instrument into a medium accuracy one. By this they achieve maximum sensitivity. Alternatively in an attempt to maximize specificity they may decide to narrow the mass tolerance window to an unreasonably small value, but this only achieves the illusion of high mass accuracy at the expense of losing true positives. For determining the appropriate size of the window, the graph of the mass error distribution is indispensable. Mass measurement error is partly made up of the systematic component caused by factors such as miscalibration, temperature drift of the calibration, or space charge. For decades internal standards, in the form of an added compound of known mass, have been used in MS to eliminate the systematic error and to obtain very high mass accuracy. For example, “peak matching methods” were used to reduce the error to a few ppm, a requirement of organic chemistry journals to prove identity of a synthesized compound. In proteomics experiments the internal standard comes for free because many of the thousand of MS and MS/MS measurements unambiguously identify peptides by their very high scores or by the fact that they are commonly occurring background peptides from keratins or trypsin autolysis products. We have used this procedure for a number of years in the open source program MSQuant (1Schulze W.X. Mann M. A novel proteomic screen for peptide-protein interactions..J. Biol. Chem. 2004; 279: 10756-10764Abstract Full Text Full Text PDF PubMed Scopus (258) Google Scholar) and find that it often improves mass accuracy severalfold, especially for TOF instruments (2Lasonder E. Ishihama Y. Andersen J.S. Vermunt A.M. Pain A. Sauerwein R.W. Eling W.M. Hall N. Waters A.P. Stunnenberg H.G. Mann M. Analysis of the Plasmodium falciparum proteome by high-accuracy mass spectrometry..Nature. 2002; 419: 537-542Crossref PubMed Scopus (558) Google Scholar). Because these internal standards are present in virtually every proteomics experiment there is no excuse for data sets with systematic error. To implement the above principles, we suggest routinely following the following steps. (i)Search the data with a permissive MMD chosen to retain essentially all correct hits.(ii)Recalibrate the mass scale by least square fitting using a few hundred hits from the highest scoring peptides and/or known contaminants.(iii)Plot the mass errors of all identified peptides. Check whether error distribution is normal. If not, remove outliers and recalibrate the mass scale using only peptides with small mass deviations and repeat (iii).(iv)Choose the MMD appropriate for the experiment from the distribution of mass errors in the plot. If the mass accuracy changed dramatically, such as more than a factor of 2 or 3, the search should be repeated with the new MMD.(v)Retain or discard peptide hits based on this MMD. These steps can readily be automated with the help of scripts implemented into proteomic data processing pipelines. Nonetheless we encourage search engine developers to implement them directly into their software. If the experiment is performed with different MS conditions or on different days, the procedure should be applied to each data set separately. The elimination of systematic error and the generation of the mass error plot can be performed completely automatically and by default. The software could let the user choose the desired trade-off between sensitivity and specificity by setting the MMD in terms of a certain multiple of the standard deviation or percentage of discarded true positives. To further help proteomics researchers embrace this quality-enhancing procedure, journals publishing proteomics results could encourage routine presentation of the plots of the distributions of mass deviations with corresponding statistical mass accuracy and the chosen MMD. High mass accuracy does imply high, at the very least isotopic, resolution and thus enables correct charge state determination and identification of the monoisotopic mass. This already means a specificity increase by a factor of 3–5 compared with unresolved isotopic distributions. Moreover as a rule of thumb the molecular mass of a peptide determined with 1 ppm mass accuracy rules out 99% of amino acid compositions possible for a given integer or nominal mass (3Zubarev R. Harkansson P. Sundqvist B. Accuracy requirements for peptide characterization by monoisotopic molecular mass measurements..Anal. Chem. 1996; 68: 4060-4063Crossref Scopus (106) Google Scholar). Thus total specificity improvement K (proportion of false possibilities filtered out) is 300–500 for ±1 ppm measurements compared with low resolution measurements. However, this estimate depends upon the number of possible alternatives, and for limited databases it can be much lower. An MMD of a few ppm is possible in routine practice (4Olsen J.V. Ong S.E. Mann M. Trypsin cleaves exclusively C-terminal to arginine and lysine residues..Mol. Cell. Proteomics. 2004; 3: 608-614Abstract Full Text Full Text PDF PubMed Scopus (864) Google Scholar, 5Olsen J.V. de Godoy L.M. Li G. Macek B. Mortensen P. Pesch R. Makarov A. Lange O. Horning S. Mann M. Parts per million mass accuracy on an Orbitrap mass spectrometer via lock mass injection into a C-trap..Mol. Cell. Proteomics. 2005; 4: 2010-2021Abstract Full Text Full Text PDF PubMed Scopus (1239) Google Scholar, 6Makarov A. Denisov E. Lange O. Horning S. Dynamic range of mass accuracy in LTQ Orbitrap hybrid mass spectrometer..J. Am. Soc. Mass Spectrom. 2006; 17: 977-982Crossref PubMed Scopus (333) Google Scholar, 7Haas W. Faherty B.K. Gerber S.A. Elias J.E. Beausoleil S.A. Bakalarski C.E. Li X. Villen J. Gygi S.P. Optimization and use of peptide mass measurement accuracy in shotgun proteomics..Mol. Cell. Proteomics. 2006; 5: 1326-1337Abstract Full Text Full Text PDF PubMed Scopus (237) Google Scholar), and consistent achievement of this value would eliminate the vast majority of false positive identifications currently reported in the literature. A very accurately measured monoisotopic molecular mass can completely specify the elemental composition of the molecule. The abundances of the isotopic peaks may be used to achieve some filter for possible composition at lower mass accuracy, but high sensitivity measurements must rely first of all on monoisotopic masses. For peptides of 1 kDa an MMD of 0.2–0.3 ppm is required (3Zubarev R. Harkansson P. Sundqvist B. Accuracy requirements for peptide characterization by monoisotopic molecular mass measurements..Anal. Chem. 1996; 68: 4060-4063Crossref Scopus (106) Google Scholar, 8Spengler B. De novo sequencing, peptide composition analysis, and composition-based sequencing: a new strategy employing accurate mass determination by Fourier transform ion cyclotron resonance mass spectrometry..J. Am. Soc. Mass. Spectrom. 2004; 15: 703-714Crossref PubMed Scopus (121) Google Scholar), and this may become available in the future. At higher masses, even this MMD is not sufficient to determine the composition, but the database is more sparsely populated with tryptic peptides in this mass range. Thus a high mass accuracy measurement may suggest a single fully tryptic peptide for a large peptide. High mass accuracy is also beneficial for finding related peaks, that is peptides that cover the same amino acid sequence but that have different mass due to modifications or missed enzymatic cleavage sites. For example, the program ModifiComb correlates peptide masses and their fragmentation spectra to a “base peptide” that is unmodified (9Savitski M.M. Nielsen M.L. Zubarev R.A. ModifiComb, a new proteomic tool for mapping substoichiometric post-translational modifications, finding novel types of modifications, and fingerprinting complex protein mixtures..Mol. Cell. Proteomics. 2006; 5: 935-948Abstract Full Text Full Text PDF PubMed Scopus (159) Google Scholar), and the program MS-Alignment groups peptides with related MS/MS spectra (10Tsur D. Tanner S. Zandi E. Bafna V. Pevzner P.A. Identification of post-translational modifications by blind search of mass spectra..Nat. Biotechnol. 2005; 23: 1562-1567Crossref PubMed Scopus (224) Google Scholar). The better the mass accuracy, the better the correlation and grouping. More generally, the better the mass accuracy, the less the need for peptide sequencing. An extremely accurate mass, perhaps combined with the elution time of the peptide, could be sufficiently characteristic of a peptide to identify it. In our experience, this is not realistic when using proteomics in a “discovery mode” but becomes quite practical in a “remeasurement mode.” We have extensively used this concept in the “protein correlation profiling” technique in which thousands of peptides are quantified across protein fractions to distinguish true members of organelles from co-migrating background proteins (11Andersen J.S. Wilkinson C.J. Mayor T. Mortensen P. Nigg E.A. Mann M. Proteomic characterization of the human centrosome by protein correlation profiling..Nature. 2003; 426: 570-574Crossref PubMed Scopus (1045) Google Scholar, 12Foster L.J. de Hoog C.L. Zhang Y. Xie X. Mootha V.K. Mann M. A mammalian organelle map by protein correlation profiling..Cell. 2006; 125: 187-199Abstract Full Text Full Text PDF PubMed Scopus (468) Google Scholar). Complete avoidance of peptide fragmentation is the basis of the “accurate mass and time” strategy (13Smith R.D. Anderson G.A. Lipton M.S. Pasa-Tolic L. Shen Y. Conrads T.P. Veenstra T.D. Udseth H.R. An accurate mass tag strategy for quantitative and high-throughput proteome measurements..Proteomics. 2002; 2: 513-523Crossref PubMed Scopus (402) Google Scholar), which will become much more reliable with the very high mass accuracy of the current generation of instruments. However, modern hybrid instruments can perform peptide fragmentation and acquisition of the tandem mass spectrum simultaneously with the acquisition of the survey mass spectrum, so it is always desirable to sequence at least a subpopulation of peptides. Several types of instruments achieve different mass accuracy in the MS versus the MS/MS mode. For example, TOF instruments have mass accuracy in the MS/MS but are often limited by ion and consequently mass accuracy for peaks and in the MS/MS mode. However, improving performance of TOF instruments them for high mass accuracy In hybrid instruments the survey spectrum is in the high resolution at the same time that the MS/MS spectrum is at high sensitivity in the low resolution of the instrument. high mass accuracy in MS/MS is clearly especially for the identification of peptides or de novo sequencing. However, this often comes at the expense of sensitivity. because the peaks are not independent but are by the limited number of possible amino acid masses and the molecular mass, the total specificity improvement K in database searches for n peaks is than In general, low resolution measurements are more for small databases and only few possible modifications, whereas high mass accuracy in the MS/MS may be required in of new or sequences and For peptides not in the high accuracy MS/MS data have the of between lysine and A further to the use of high mass accuracy data is the fact that database do not yet make full use of accurate MS/MS For example, there is a in mass accuracy which the program to higher to higher accuracy the multiple Fourier with become with high resolution Thus proteomics researchers do not obtain higher scores from higher resolution and as a they are to data in high resolution even when this is easily we an to peptide identification to high accuracy The current use of mass accuracy is still quite We usually set a for the MMD and completely that is within the whereas we completely peaks with a very signal and signal to noise are the same as peaks to noise We propose that mass accuracy should be in a more by a mass deviation defined peptide peaks must be much more in terms of MMD than peaks to the of each peptide should have mass This could be based on repeat measurement an J.V. de Godoy L.M. Li G. Macek B. Mortensen P. Pesch R. Makarov A. Lange O. Horning S. Mann M. Parts per million mass accuracy on an Orbitrap mass spectrometer via lock mass injection into a C-trap..Mol. Cell. Proteomics. 2005; 4: 2010-2021Abstract Full Text Full Text PDF PubMed Scopus (1239) Google this also has the that the mass from the it can be determined best is given compared with the spectrum in which the was for and it is often very peptide molecular masses should be determined from the available data set or data sets and not only from a single in a single mass In the mass deviation should be of the database search with for peptides measured with lower mass deviation. Such a has already been into at least one peptide identification A. E. R. statistical model to estimate the accuracy of peptide identifications made by MS/MS and database Chem. 2002; PubMed Scopus Google Scholar). The of such in to of the mass accuracy, is that they can easily such as the shape of the isotopic distribution, the of deviation of the time of the peptide from the expected value, and so In MS-based proteomics can be extremely and It is a of our field that our data on an objective and parameter of each mass, and we should on this This is especially true now when we can determine masses to many to the of determining the elemental This high accuracy is in to many other of which have large errors and of It potentially MS-based proteomics one of the most “digital” and of the at least as peptide It is up to to make the most of this by the best possible of mass accuracy. We V. for and for the figure.
No takes yet. Share an insight, caveat, or question.
Zubarev et al. (2006) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: