Targeted mass spectrometry is an essential tool for detecting quantitative changes in low abundant proteins throughout the proteome. Although selected reaction monitoring (SRM) is the preferred method for quantifying peptides in complex samples, the process of designing SRM assays is laborious. Peptides have widely varying signal responses dictated by sequence-specific physiochemical properties; one major challenge is in selecting representative peptides to target as a proxy for protein abundance. Here we present PREGO, a software tool that predicts high-responding peptides for SRM experiments. PREGO predicts peptide responses with an artificial neural network trained using 11 minimally redundant, maximally relevant properties. Crucial to its success, PREGO is trained using fragment ion intensities of equimolar synthetic peptides extracted from data independent acquisition experiments. Because of similarities in instrumentation and the nature of data collection, relative peptide responses from data independent acquisition experiments are a suitable substitute for SRM experiments because they both make quantitative measurements from integrated fragment ion chromatograms. Using an SRM experiment containing 12,973 peptides from 724 synthetic proteins, PREGO exhibits a 40–85% improvement over previously published approaches at selecting high-responding peptides. These results also represent a dramatic improvement over the rules-based peptide selection approaches commonly used in the literature. Targeted mass spectrometry is an essential tool for detecting quantitative changes in low abundant proteins throughout the proteome. Although selected reaction monitoring (SRM) is the preferred method for quantifying peptides in complex samples, the process of designing SRM assays is laborious. Peptides have widely varying signal responses dictated by sequence-specific physiochemical properties; one major challenge is in selecting representative peptides to target as a proxy for protein abundance. Here we present PREGO, a software tool that predicts high-responding peptides for SRM experiments. PREGO predicts peptide responses with an artificial neural network trained using 11 minimally redundant, maximally relevant properties. Crucial to its success, PREGO is trained using fragment ion intensities of equimolar synthetic peptides extracted from data independent acquisition experiments. Because of similarities in instrumentation and the nature of data collection, relative peptide responses from data independent acquisition experiments are a suitable substitute for SRM experiments because they both make quantitative measurements from integrated fragment ion chromatograms. Using an SRM experiment containing 12,973 peptides from 724 synthetic proteins, PREGO exhibits a 40–85% improvement over previously published approaches at selecting high-responding peptides. These results also represent a dramatic improvement over the rules-based peptide selection approaches commonly used in the literature. Targeted proteomics using selected reaction monitoring (SRM)1 and parallel reaction monitoring (PRM) is increasingly becoming the gold-standard method for peptide quantitation within complex biological matrices (1.Marx V. Targeted proteomics.Nat. Methods. 2013; 10: 19-22Crossref PubMed Scopus (147) Google Scholar, 2.Liebler D.C. Zimmerman L.J. Targeted quantitation of proteins by mass spectrometry.Biochemistry. 2013; 52: 3797-3806Crossref PubMed Scopus (241) Google Scholar). By focusing on monitoring only a handful of transitions (associated precursor and fragment ions) for targeted peptides, SRM experiments filter out background signals, which in turn increases the signal to noise ratio. SRM experiments are almost exclusively performed on triple-quadrupole instruments. These instruments can isolate single transitions as an ion beam and measure that beam with extremely sensitive ion-striking detectors. As a result, SRM experiments generally exhibit significantly more accurate quantitation when compared with similarly powered discovery based proteomics experiments, and frequently benefit from a much wider linear range of quantitation (3.Picotti P. Aebersold R. Selected reaction monitoring-based proteomics: workflows, potential, pitfalls, and future directions.Nat. Methods. 2012; 9: 555-566Crossref PubMed Scopus (996) Google Scholar). SRM experiments often require less fractionation and can be run in shorter time on less expensive instrumentation. These factors allow researchers to greatly scale up the number of samples they can run, which in turn increases the power of their experiment. However, the process of developing an effective SRM assay is often cumbersome, as subtle differences in peptide sequence can have a profound impact on the physiochemical properties and subsquent SRM responses of a peptide. To successfully develop an SRM assay for a protein of interest, unique peptide sequences must be chosen that also produce a high SRM signal (e.g. high-responding peptides). Once identified, these high-responding peptides are often synthesized or purchased, and independently analyzed to determine the most sensitive transition pairs. Finally, the selected peptide and transition pairs must be tested in complex mixtures to screen for transitions with chemical noise interference and to validate the sensitivity of the assay within a particular sample matrix. Peptides and transitions that survive this lengthy screening process can then undergo absolute quantitation by calibrating the signal intensity against standards of known quantity. Although experimental methods have been developed to empirically determine a set of best responding peptides (4.Stergachis A.B. MacLean B. Lee K. Stamatoyannopoulos J.A. MacCoss M.J. Rapid empirical discovery of optimal peptides for targeted proteomics.Nat. Methods. 2011; 8: 1041-1043Crossref PubMed Scopus (90) Google Scholar), these strategies can be time consuming and require analytical standards, which are currently unavailable for all proteins. More often than not, representative peptides are essentially chosen at random, using only a small number of criteria, such as having a reasonable length for detection in the mass spectrometer, a lack of methionine, and a preference for peptides containing proline (5.Bereman M.S. MacLean B. Tomazela D.M. Liebler D.C. MacCoss M.J. The development of selected reaction monitoring methods for targeted proteomics via empirical refinement.Proteomics. 2012; 12: 1134-1141Crossref PubMed Scopus (83) Google Scholar). It is not uncommon for SRM assays to fail at the final validation steps simply because the peptides chosen in the first assay creation step happened to be unexpectedly poor responding peptides. In an effort to speed up the process of generating robust assays, several groups (6.Mallick P. Schirle M. Chen S.S. Flory M.R. Lee H. Martin D. Ranish J. Raught B. Schmitt R. Werner T. Kuster B. Aebersold R. Computational prediction of proteotypic peptides for quantitative proteomics.Nat. Biotechnol. 2007; 25: 125-131Crossref PubMed Scopus (570) Google Scholar, 7.Fusaro V.A. Mani D.R. Mesirov J.P. Carr S.A. Prediction of high-responding peptides for targeted protein assays by mass spectrometry.Nat. Biotechnol. 2009; 27: 190-198Crossref PubMed Scopus (235) Google Scholar, 8.Eyers C.E. Lawless C. Wedge D.C. Lau K.W. Gaskell S.J. Hubbard S.J. CONSeQuence: Prediction of reference peptides for absolute quantitative proteomics using consensus machine learning approaches.Mol. Cell. Proteomics. 2011; 10M110.003384Abstract Full Text Full Text PDF PubMed Scopus (101) Google Scholar, 9.Muntel J. Boswell S.A. Tang S. Ahmed S. Wapinski I. Foley G. Steen H. Springer M. Abundance-based classifier for the prediction of mass spectrometric peptide detectability upon enrichment.Mol. Cell. Proteomics. 2015; 14: 430-440Abstract Full Text Full Text PDF PubMed Scopus (18) Google Scholar) have designed approaches to predict sets of proteotypic peptides using machine-learning algorithms. Proteotypic peptides are peptides commonly identified in shotgun proteomics experiments for a variety of reasons including high signal, low interference, and search engine compatible fragmentation. Enhanced Signature Peptide (ESP) Predictor (7.Fusaro V.A. Mani D.R. Mesirov J.P. Carr S.A. Prediction of high-responding peptides for targeted protein assays by mass spectrometry.Nat. Biotechnol. 2009; 27: 190-198Crossref PubMed Scopus (235) Google Scholar) was the first successful modification of this prediction approach to use proteotypic peptides as a proxy for high-responding peptides for SRM-based quantitation. In brief, Fusaro et al. built a training data set from data-dependent acquired (DDA) yeast peptides and a proxy for their response was quantitated using extracted precursor ion chromatograms (XICs). The authors calculated 550 physiochemical properties for each peptide based on sequence alone and built a random forest classifier to differentiate between the high and low response groups. Other peptide prediction tools follow the same general methodology for developing training data sets. CONSeQuence (8.Eyers C.E. Lawless C. Wedge D.C. Lau K.W. Gaskell S.J. Hubbard S.J. CONSeQuence: Prediction of reference peptides for absolute quantitative proteomics using consensus machine learning approaches.Mol. Cell. Proteomics. 2011; 10M110.003384Abstract Full Text Full Text PDF PubMed Scopus (101) Google Scholar) applies several machine learning strategies and a pared down list of 50 distinct peptide properties. Alternately, Peptide Prediction with Abundance (9.Muntel J. Boswell S.A. Tang S. Ahmed S. Wapinski I. Foley G. Steen H. Springer M. Abundance-based classifier for the prediction of mass spectrometric peptide detectability upon enrichment.Mol. Cell. Proteomics. 2015; 14: 430-440Abstract Full Text Full Text PDF PubMed Scopus (18) Google Scholar) (PPA) uses a back-propagation neural network (10.Rumelhart D.E. Hinton G.E. by Scopus Google Scholar) trained with distinct peptide properties selected from The authors of CONSeQuence and that their approaches the Predictor on a variety of data sets. As with most machine the of the training set to data is to the of the prediction Although intensities extracted from data can be for high-responding peptides Tomazela D.M. B. B. G. S. M.J. the development of targeted SRM using data from shotgun proteomics to method 2009; 8: PubMed Scopus Google Scholar, J.A. V. C. C. the tool for designing reaction monitoring Cell. Proteomics. 2009; 8: Full Text Full Text PDF PubMed Scopus Google Scholar), several factors make less than for to SRM and experiments. In peptides must be identified and for targeted of transitions can by fragment that are to with search By training data sets on precursor intensities alone these tools the that targeted assays use fragment for that training sets from fragment intensities MacLean B. R. MacCoss M.J. peptide using data independent acquisition and 2015; 10: PubMed Scopus Google Scholar) produce machine-learning tools that are more effective at peptides that produce than proteotypic peptides. The use of proteins in training sets The in peptide intensities is by in protein abundance. peptide intensities to can the on varying protein at the of the training set with proteins that high-responding peptides. to this by training with B. D. G. J. J. Chen M. of 2011; PubMed Scopus Google Scholar) for peptides from that a training set from equimolar synthetic peptides most of from the training to a more of peptides or as a from Peptide peptides are of proteins that have been in shotgun of peptide selection a small peptides that can be with fractionation was to of the peptides. Peptides acquired with all to In the training peptides are representative of peptides with one the training data set not peptides with a of the peptide of each was in of and The was for at for of the was in of for a which was down to to a sample for or of the was a with The sample was and using of a The was with the analytical The analytical was a with a using a The analytical was with of The analytical was to a and Peptides of the at a of using a of in by in over Peptides by and a mass acquired using one of acquisition data-dependent acquisition (DDA) or acquisition The method an with target and time 50 up to from the most in the The have target time with an intensity an or The time was with of targeted and the set to was acquired with target and time the each with a target time with the set to The and the range from The of of was as and The of and acquisition and was throughout the The data was using against a containing the peptide to with the been using M.R. MacCoss M.J. speed data and of shotgun proteomics using mass 2007; Scopus Google Scholar) and M.R. MacLean B. MacCoss M.J. of search strategies for high precursor mass 9: PubMed Scopus Google Scholar) to more accurate precursor based on of and a The with J. MacCoss M.J. learning for peptide from shotgun proteomics Methods. 2007; PubMed Scopus Google Scholar) to to peptide and peptide G.E. MacCoss M.J. of peptide from proteomics experiments using PubMed Scopus Google Scholar) was used to the a containing with The is extremely because the is simply used as an for of the The data analyzed using the B. Tomazela D.M. M. B. R. Liebler D.C. MacCoss M.J. an for and targeted proteomics PubMed Scopus Google Scholar) software In chromatograms extracted for the precursor of each peptide that within the analyzed each peptide chromatograms extracted for the and precursor from the and chromatograms for the to ion extracted from the The for each peptide precursor selected and integrated in each of the data sets The time of from the data on the data to in selecting the the mass in of the of the precursor to the and in the of the of the extracted fragment ion chromatograms from the data to in the used to that the was In the of was a all of these this was not the the peptide precursor was in a of peptides interference also The data in et al. was used as a data SRM training validation data set was using the in et al. for proteins from the Scholar) synthesized in using the in protein In a was proteins from a proteins using and to proteins with for at and with for at then with of for at on a analytical with The analytical was to a and Peptides the at a of using in and in this linear Peptides are by and a peptides of length to for each protein analyzed using the software Peptide fragment chromatograms for the to ion extracted from the data and of the proteins used for training validation to against over The proteins exclusively for a data set and used only training was Peptide responses for peptides in the et al. SRM data set using and at was used using the mass from to and peptide length of The artificial neural network and linear machine of CONSeQuence at run independent of the consensus The consensus was not used because only which to against the Predictor at is Peptide response factors within proteins on by over of between the and responding peptides. et al. previously an experimental method for the best responding peptides to proteins in targeted experiments. method was by over factors in and generating SRM assays for all to from peptide. Because of in proteins in this experiment not at the same However, all peptides within a protein to be present at equimolar and using this the authors to determine which peptides the best SRM transitions for in In this we use the et al. data set as an independent set to validate of this data set for that was acquired only precursor peptides against high peptides and and that analyzed fragment to only that the of the scale of this data set these The et al. data set an for the in peptide the range of peptide transition responses in the et al. SRM data Although the range of peptide responses within a protein was of proteins response of up to or of for a with an range of of is in of responses the for a robust for peptides to In this we the et al. data set containing 12,973 peptides from 724 proteins a of peptides protein and a of to approach for peptide responses for and data sets that are to are for effective machine However, an targeted data set of equimolar peptides for training a peptide response prediction is extremely time consuming as require SRM experiments to for all transition for peptide. have developed a for generating SRM and training sets using experiments acquired on a using fragmentation. the of a training data the that all sequence are to the most we used to which is to used in most SRM experiments B. S. SRM assay a between ion and peptide 2011; 10: PubMed Scopus Google Scholar). the training set from the most fragment intensity for each of peptide by used because can to in both response and the in instruments and to at the fragment ion is frequently one of the most in the we list of fragment to for each we the fragment by mass in both this we the fragment for each peptide as a proxy for the transition Because peptide from pairs of at and we to use the of quantitative to peptides on this we peptides from the peptides that than or than from In peptides to in between and that their intensities Peptides because their intensities also these peptides, we the of the pairs of to be the ratio. the intensity for each peptide as the of the intensities from the acquisition and the intensities and the peptides with the and to for peptides with in a final training data set of peptides, which are in these peptides are in Finally, we the peptides in the training set based on these fragment ion intensities and the to be between and each peptide sequence we calculated 550 physiochemical properties used by the of which from the S. M. PubMed Google Scholar). out that one of is that used in this in proteomics are the of the properties are the for these properties to be between and selected physiochemical properties using a C. selection from PubMed Scopus Google Scholar, H. C. selection based on of and 27: PubMed Scopus Google Scholar). each we calculated the of peptides with the from their peptide The with the was selected as a and all properties that with that at an absolute of are process is using the properties all properties that have to the intensity are selected or The 11 most relevant physiochemical properties. These properties and their to the training intensities are in I. As the the most representative of several the properties are less than their Peptides with with high transition intensities in training by and relevant physiochemical peptide properties selected from a of 550 properties based on their with the intensity in the training data are based on the absolute of the which is an of their for As each was with properties with peptide and relative preference at D.C. for at the of PubMed Scopus Google of K. K. T. of on of the in a of proteins at a unique of PubMed Scopus Google in of of of from of Scopus Google of of empirical methods for the prediction of protein PubMed Scopus Google from data set J. of protein their and in protein PubMed Scopus Google in of of proteins H. K. The is between the and in PubMed Scopus Google of in J. P. P. on the of prediction J. Peptide Scopus Google relative of M. G. of the and of the Scopus Google of in proteins of S. K. between and PubMed Scopus Google Peptide properties selected from a of 550 properties based on their with the intensity in the training data are based on the absolute of the which is an of their for As each was with Peptide properties with peptide and in a The final training set of the and the of peptides to between high and low responding peptides, the was the intensity a back-propagation neural network with 11 to the 11 relevant physiochemical in a single and a single the neural network for a learning and trained to a of produce a between and the of the of using the neural network the PREGO was PREGO in an effort to that and is to the for of the PREGO is in are to make when a machine learning As with and we to an artificial neural network because to than (e.g. on 2007; Scholar). However, the machine approach to back-propagation is random in artificial neural to often on than we trained and using proteins selected from an SRM data set the et al. experiment. selected the best that the of the that compared the number of peptides protein the number of proteins at one high-responding peptide was each peptides high they a single most fragment ion for each peptide in the of peptides from that approach also a against because we trained using data and the training with SRM data acquired in a PREGO using the et al. data which experimental SRM transition responses acquired for almost peptides in over proteins. with we this data set to using the only the single most fragment ion to the used the of PREGO for a representative protein in this data a of when compared with the experimental intensity the of the all proteins in the data set Although is in PREGO are generally high in of high-responding peptides, and low with less peptides. the range of PREGO for a variety of proteins that with from to in all proteins in the et al. data the of PREGO for peptides at in all of the proteins, the the and the the the in is at each However, the in as that PREGO is to differentiate peptide responses in SRM experiments. a similarly for on the same set of proteins. Although is a in the high to peptides at all The of the that is more to low to low responding peptides. of these low responding peptides from the of and increases the for a high-responding peptide. CONSeQuence using both the artificial neural network and the are in and In this data CONSeQuence a in with responding the in the major Although is that response prediction with experimental peptide these be used to peptides to a protein in the that at one a The approaches not the responding peptide to be effective they must be to at one responding peptide in a handful of the we selected peptides for at one of peptides high high response as in the of peptides for each protein by these criteria, on PREGO a high-responding peptide of the time on the first peptides protein are then at one is a high of the and on selecting peptides a high of the each of these PREGO high to more often than the best As a for selecting peptides at However, peptides to SRM and assays by several selection and the peptides that built a to the et al. that for produce and for can be can be in the can to and in the can also The rules-based is a of all of the in a based this than the of the relative improvement of PREGO and the trained approaches over the based of the trained approaches over the based approach when only the peptide. However, is that only a single peptide protein for targeted As one more peptides at random, is an that at one is a high-responding peptide and that increasingly to a is that when or more peptides from the et al. data simply using the et al. essentially to the and CONSeQuence PREGO, on the to over the based approach when a number of peptides for targeted results using the proteins from the SRM experiment It is to that in the of peptides for SRM and assays of is The factors that determine peptide response are and are in number and the of generating targeted assays by selecting peptides at random using of the in et al. over these is the that peptide response prediction be compared training sets and machine learning and both CONSeQuence produce essentially that these software tools than selecting peptides for SRM assays, not significantly than using a rules-based random approach for peptide response that be a for SRM response based on peptide responses in data sets. results that the PREGO a dramatic improvement over these methods for SRM Although the we we that the of from training data set In we that training from data sets using the to more represent data acquisition strategies by SRM instruments. In to more predict transition response from peptide of that precursor intensities with fragment that is an of between and precursor intensities which that training using transition responses to be more accurate than training from improvement is that PREGO robust by the trained artificial neural network with SRM As mass and can have a profound on peptide training using of data from is also that the of and CONSeQuence be by of data acquisition in the data set was to only precursor and peptide response was using only the single most fragment ion from each peptide. These represent commonly in SRM assays and the training of PREGO not in or Peptide response prediction can also be used to search that this approach to data sets can benefit from sensitivity using an data However, by peptide for all proteins in a the approach from a significantly discovery that must be for using which sensitivity of for PREGO can down the search by first only a handful of high-responding peptides search engine then only to for peptides are make one major in the of training we that peptides in are essentially at equimolar make this because developing a training set from peptides be that these peptides are between and that is less than in their this is the of in present for each we to use biological samples, such as with the or CONSeQuence also that the of the that high peptides in each protein produce high fragment ion intensities in using peptides. the training using the single most fragment ion for each peptide PREGO peptides with the most fragment ion by from the most fragment ion by SRM can be to produce the most and to on a varying in are also not for with synthetic peptides. be an from the of machine learning in that training are on peptide sequences that produce than by to of at the same The of are to in this experiment because the et al. SRM data set only peptides with However, can be a when particular of peptides, for In the future of training or for It is to that PREGO than is in the for each peptide. is because peptide transition response is the of complex only of which can be using physiochemical properties. The for peptide response as experimental from synthetic proteins. The of PREGO is in experimental data from is or to for improvement with future prediction methods to use more training data sets and more complex properties for proteomics methods that and and present a PREGO, to high-responding peptides to in generating SRM and approach uses experimental data of equimolar synthetic peptides to an artificial neural network using 11 selected with a have software using a SRM data set peptide from over proteins.
No takes yet. Share an insight, caveat, or question.
Searle et al. (2015) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: