Key points are not available for this paper at this time.
Improvements in mass spectrometry (MS)-based peptide sequencing provide a new opportunity to determine whether polymorphisms, mutations, and splice variants identified in cancer cells are translated. Herein, we apply a proteogenomic data integration tool (QUILTS) to illustrate protein variant discovery using whole genome, whole transcriptome, and global proteome datasets generated from a pair of luminal and basal-like breast-cancer-patient-derived xenografts (PDX). The sensitivity of proteogenomic analysis for singe nucleotide variant (SNV) expression and novel splice junction (NSJ) detection was probed using multiple MS/MS sample process replicates defined here as an independent tandem MS experiment using identical sample material. Despite analysis of over 30 sample process replicates, only about 10% of SNVs (somatic and germline) detected by both DNA and RNA sequencing were observed as peptides. An even smaller proportion of peptides corresponding to NSJ observed by RNA sequencing were detected (<0.1%). Peptides mapping to DNA-detected SNVs without a detectable mRNA transcript were also observed, suggesting that transcriptome coverage was incomplete (∼80%). In contrast to germline variants, somatic variants were less likely to be detected at the peptide level in the basal-like tumor than in the luminal tumor, raising the possibility of differential translation or protein degradation effects. In conclusion, this large-scale proteogenomic integration allowed us to determine the degree to which mutations are translated and identify gaps in sequence coverage, thereby benchmarking current technology and progress toward whole cancer proteome and transcriptome analysis. Improvements in mass spectrometry (MS)-based peptide sequencing provide a new opportunity to determine whether polymorphisms, mutations, and splice variants identified in cancer cells are translated. Herein, we apply a proteogenomic data integration tool (QUILTS) to illustrate protein variant discovery using whole genome, whole transcriptome, and global proteome datasets generated from a pair of luminal and basal-like breast-cancer-patient-derived xenografts (PDX). The sensitivity of proteogenomic analysis for singe nucleotide variant (SNV) expression and novel splice junction (NSJ) detection was probed using multiple MS/MS sample process replicates defined here as an independent tandem MS experiment using identical sample material. Despite analysis of over 30 sample process replicates, only about 10% of SNVs (somatic and germline) detected by both DNA and RNA sequencing were observed as peptides. An even smaller proportion of peptides corresponding to NSJ observed by RNA sequencing were detected (<0.1%). Peptides mapping to DNA-detected SNVs without a detectable mRNA transcript were also observed, suggesting that transcriptome coverage was incomplete (∼80%). In contrast to germline variants, somatic variants were less likely to be detected at the peptide level in the basal-like tumor than in the luminal tumor, raising the possibility of differential translation or protein degradation effects. In conclusion, this large-scale proteogenomic integration allowed us to determine the degree to which mutations are translated and identify gaps in sequence coverage, thereby benchmarking current technology and progress toward whole cancer proteome and transcriptome analysis. Massively parallel sequencing (MPS) 1The abbreviations used are:iTRAQIsobaric Tags for Relative and Absolute QuantitationLFlabel freeMPSmassively parallel sequencingNGSnext generation sequencingNSJnovel splice junctionPDXpatient-derived xenograftPSMpeptide spectral matchQUILTSQuantitative Integrated Library of Translated SNPs/SplicingSNVsingle nucleotide variant. of cancer genomes has demonstrated enormous complexity, and it is often unclear which somatic mutations drive tumor biology and which are nonfunctional passenger mutations that passively accumulate. RNA sequencing is frequently used to determine which nucleotide variants are transcribed and therefore have the potential for biological function. However, many mutations detected at the DNA level are not observed at the mRNA level, and their observation is dependent upon expression of the stability of the mRNA (1.Cirulli E.T. Singh A. Shianna K.V. Ge D. Smith J.P. Maia J.M. Heinzen E.L. Goedert J.J. Goldstein D.B. Screening the human exome: A comparison of whole genome and whole transcriptome sequencing.Genome Biol. 2010; 11: R57Crossref PubMed Scopus (104) Google Scholar). Mutation detection at the peptide level clearly increases the confidence that any given variant is a potential biological driver, and by assessing peptide levels across all of an individual's polymorphisms, an independent assessment of transcriptome coverage can be obtained. Isobaric Tags for Relative and Absolute Quantitation label free massively parallel sequencing next generation sequencing novel splice junction patient-derived xenograft peptide spectral match Quantitative Integrated Library of Translated SNPs/Splicing single nucleotide variant. Integrated proteogenomic methods that combine MPS analysis and proteomics are of particular importance for identifying novel peptides resulting from somatic mutations or inherited polymorphisms. The identification of peptide sequences by mass spectrometry (MS) relies heavily on the quality of the protein sequence database. Use of databases with missing peptide sequences will fail to identify the corresponding peptides within the proteomic data; however, addressing this by including large numbers of irrelevant sequences in the search will decrease sensitivity. Therefore, it is essential that data acquired through MPS are used to create tumor-specific databases, incorporating the possibility of variant proteins arising through somatic mutation, inherited polymorphisms, alternatively spliced isoforms, and novel expression. The goal of this study was to analyze the flow of information though the central dogma of biology in an unbiased and comprehensive way to profoundly understand the aberrant information flux that underlies all cancer biology (2.Wang X. Slebos R.J. Wang D. Halvey P.J. Tabb D.L. Liebler D.C. Zhang B. Protein identification using customized protein sequence databases derived from RNA-Seq data.J. Proteome Res. 2012; 11: 1009-1017Crossref PubMed Scopus (130) Google Scholar, 3.Castellana N.E. Shen Z. He Y. Walley J.W. Cassidy C.J. Briggs S.P. Bafna V. An automated proteogenomic method utilizes mass spectrometry to reveal novel genes in Zea mays.Mol. Cell. Proteomics. 2013; 13: 157-167Abstract Full Text Full Text PDF PubMed Scopus (65) Google Scholar, 4.Li J. Su Z. Ma Z.Q. Slebos R.J. Halvey P. Tabb D.L. Liebler D.C. Pao W. Zhang B. A bioinformatics workflow for variant peptide detection in shotgun proteomics.Mol. Cell. Proteomics. 2011; 10 (M110.006536)Abstract Full Text Full Text PDF Scopus (84) Google Scholar, 5.Sheynkman G.M. Shortreed M.R. Frey B.L. Smith L.M. Discovery and mass spectrometric analysis of novel splice-junction peptides using RNA-Seq.Mol. Cell. Proteomics. 2013; Full Text Full Text PDF PubMed Scopus Google Scholar). Integrated proteogenomic have in including and human (2.Wang X. Slebos R.J. Wang D. Halvey P.J. Tabb D.L. Liebler D.C. Zhang B. Protein identification using customized protein sequence databases derived from RNA-Seq data.J. Proteome Res. 2012; 11: 1009-1017Crossref PubMed Scopus (130) Google Scholar, 4.Li J. Su Z. Ma Z.Q. Slebos R.J. Halvey P. Tabb D.L. Liebler D.C. Pao W. Zhang B. A bioinformatics workflow for variant peptide detection in shotgun proteomics.Mol. Cell. Proteomics. 2011; 10 (M110.006536)Abstract Full Text Full Text PDF Scopus (84) Google Scholar, 5.Sheynkman G.M. Shortreed M.R. Frey B.L. Smith L.M. Discovery and mass spectrometric analysis of novel splice-junction peptides using RNA-Seq.Mol. Cell. Proteomics. 2013; Full Text Full Text PDF PubMed Scopus Google Scholar, splice variants, a new of protein cancer in cancer and cancer with biology 2010; PubMed Google Scholar, J. of from for transcript and protein 2012; PubMed Scopus Google and in for N.E. Shen Z. He Y. Walley J.W. Cassidy C.J. Briggs S.P. Bafna V. An automated proteogenomic method utilizes mass spectrometry to reveal novel genes in Zea mays.Mol. Cell. Proteomics. 2013; 13: 157-167Abstract Full Text Full Text PDF PubMed Scopus (65) Google Scholar, Wang Y. D. X. D. A. P. A. A. W. and of transcriptome by parallel sequencing and shotgun Res. 2011; PubMed Scopus Google Scholar, He Y. Bafna V. from large data.J. Proteome Res. 2013; 13: PubMed Scopus Google Scholar, X. X. J. Y. The discovery of novel in genome on mass spectrometry 2011; PubMed Scopus Google Scholar). DNA databases in MS identification are a we a of the proteomic and sensitivity to a comprehensive or germline variant peptide we used patient-derived xenograft cancer that for and proteomic to determine the and of MS/MS for novel protein identification in cancer xenograft from and cancer D. He X. Z. J. of cancer on PubMed Scopus Google A. D. of human PubMed Scopus Google were from the and of The were in as J.W. J. D.C. J. Zhang Y. Smith W. Shen D. J.M. P.J. G.M. in a basal-like cancer and 2010; PubMed Scopus Google Shen D. J. W. A. He X. J. D. J. Y. A. Zhang J. D. Ma W. Wang A. variants by of 2013; Full Text Full Text PDF PubMed Scopus Google and in with and and by the at have expression and proteomic Shen D. J. W. A. He X. J. D. J. Y. A. Zhang J. D. Ma W. Wang A. variants by of 2013; Full Text Full Text PDF PubMed Scopus Google that are to their biology and were and to a with as P. J.W. R.J. P. K.V. D. Liebler D. Smith in and in not global protein Cell. Proteomics. 13: Full Text Full Text PDF PubMed Scopus Google Scholar, B. Wang J. Wang X. J. Z. Wang Wang P. Tabb D.L. R.J. Slebos R.J. Liebler D.C. of human and PubMed Scopus Google Scholar). from were by at with by in a The tumor were in on and at A of was in to that be and multiple tumor were and in a using to the tumor was to an on and the was with a in The was will to a were on to in a and data have Shen D. J. W. A. He X. J. D. J. Y. A. Zhang J. D. Ma W. Wang A. variants by of 2013; Full Text Full Text PDF PubMed Scopus Google Scholar). using the data were by an variant D.C. G.M. detection in massively parallel sequencing of and PubMed Scopus Google A. A. A. D. The A for DNA sequencing Res. 2010; PubMed Scopus Google and Z. A to of large and from PubMed Scopus Google to and sensitivity. the RNA-Seq were by from both RNA-Seq were to the human genome using with and to junction and D. An for discovery of novel Biol. 2011; PubMed Scopus Google and detection D.C. G.M. detection in massively parallel sequencing of and PubMed Scopus Google Scholar, A. A. A. D. The A for DNA sequencing Res. 2010; PubMed Scopus Google were used to potential and were used to the pair for on using in were in and in a with derived from the of the tumor The resulting was to a using a and were on to sample were and to Peptides were using an or and by and The analysis was by and the analysis was at The data are through the is a that can in to for a RNA-Seq somatic variants and germline and a all The of is a protein sequence using or as a for the proteome and and are the for and of variant and nucleotide from the variant for In both somatic and germline variants are tumor-specific variants are by all germline variants from the somatic variant on from or sequences of are and variant were on The sequences are translated to proteins in a single and as a and to single were for and within In this even variants were in the variant to for by mass spectrometric analysis. The junction of for splice or only novel junction which are novel and and novel The translation for in protein is in for peptides are the sequences of the alternatively spliced proteins as in In this with at RNA were in by a translation Protein with than are in the protein database. was used to MS data to for data analysis. spectral was using the tumor-specific databases with the search using of and mass of for data and for of and on peptide and potential of of and of and for The databases a of and peptides for and databases a of from of by the global proteome in to sequences for The search was a and human and luminal variant and NSJ peptides. were for using a Discovery with that were through of all protein sequences and in protein identification by sequence databases using mass spectrometry PubMed Scopus Google Scholar). The was as by to peptides identified by tandem mass spectrometry using Proteome Res. PubMed Scopus Google and was for at a of at the peptide The was for and peptide and 11: PubMed Scopus Google Scholar). identified variant and junction peptides are The was also the databases, to and any that in this search a variant peptide not detected in the data was from the analysis. peptides with than were The were on the of the in the mass that were by of the peptide sequence and only peptides with than of the to the sequence with than of the were Peptides were also not to have gaps of than peptides with in a without of were peptides with or without were In the of gaps a peptide the peptide was The for the on was the observation is observed for of a peptide not for the the identification is In often for that the search has the and the sequence is missing from the that was spectrometry and by analysis of and luminal was across using or The of MS/MS of sample process replicates, and the single nucleotide variant (SNV) and novel junction (NSJ) peptides are given for used for in a new Proteome analysis of and luminal was across using or The of MS/MS of sample process replicates, and the single nucleotide variant (SNV) and novel junction (NSJ) peptides are given for used for the quality of the variant peptide the of the of MS/MS by the identified the peptide the and were with the for all identified peptides in the protein sequence and the of the variant and and and peptides in the MS/MS by the identified peptide and MS/MS were was observed, we that the for the variant peptides can be using the peptide for the NSJ peptides the quality was by the of the of MS/MS by the identified the peptide the and to the for all identified peptides in the protein sequence the of the NSJ and and and peptides in the MS/MS by the identified peptide and MS/MS were was observed the NSJ peptides and the that the quality of NSJ peptides is to that of the peptides. In was observed for germline and somatic variant peptides or for variant peptides with or without mRNA the peptide are in The are with the and and The for of is also by the peptide sequence with with to the corresponding variant and novel peptide a of and 30 sample process replicates tandem MS using identical sample were used to identify peptides. Peptides with at peptide spectral match were as a is a of the that the of gaps in the peptide sequence coverage than and a of gaps over the peptides the or were peptides that only to variant or novel junction peptides. variant and NSJ peptides were with the protein sequence to any peptides that were by the all novel peptides were the translation and the protein P. A. A. D. Zhang Y. A. The on human Res. PubMed Scopus Google and all were in the and Peptides with were in all analysis in the of variant and novel junction peptide for which were of the process replicates, the of variant or novel peptides identified was used for analysis. was used to determine the peptides identified and peptides in were the and the of in to determine the variant identified for proteogenomic by and 11: PubMed Scopus Google the protein sequence databases here are at for all novel splice and single nucleotide variant peptides are tumor were generated from as J.W. A. of and human in A. PubMed Scopus Google Scholar). derived from a luminal tumor, and the from a basal-like tumor, were in this of which the whole genome sequences and RNA-Seq analysis have J.W. J. D.C. J. Zhang Y. Smith W. Shen D. J.M. P.J. G.M. in a basal-like cancer and 2010; PubMed Scopus Google Scholar, Shen D. J. W. A. He X. J. D. J. Y. A. Zhang J. D. Ma W. Wang A. variants by of 2013; Full Text Full Text PDF PubMed Scopus Google Scholar). In the coverage of was for tumor, and RNA-Seq data of were for proteome expression was acquired by independent MS/MS sample process replicates analysis the luminal and both of which were by in the and independent MS/MS sample were across using and using methods The of variant and novel junction peptides identified by on MS and A of human peptides and proteins were identified across all the proteomic coverage with current tumor-specific protein databases and corresponding proteomics tandem MS data were used to identify novel protein in cancer and to determine the of proteomic analysis to comprehensive variant peptide identification through whole genome sequencing the of in protein sequence the of tumor-specific peptides that can be identified by tandem MS (2.Wang X. Slebos R.J. Wang D. Halvey P.J. Tabb D.L. Liebler D.C. Zhang B. Protein identification using customized protein sequence databases derived from RNA-Seq data.J. Proteome Res. 2012; 11: 1009-1017Crossref PubMed Scopus (130) Google Scholar, 5.Sheynkman G.M. Shortreed M.R. Frey B.L. Smith L.M. Discovery and mass spectrometric analysis of novel splice-junction peptides using RNA-Seq.Mol. Cell. Proteomics. 2013; Full Text Full Text PDF PubMed Scopus Google Scholar, He Y. Bafna V. from large data.J. Proteome Res. 2013; 13: PubMed Scopus Google Scholar). MS identification methods on databases, as or for peptide databases not novel peptide identification to in the somatic or germline In to identify tumor-specific variant we used the proteogenomic integration tool Quantitative Integrated Library of Translated SNPs/Splicing for cancer proteome analysis protein variant and and junction to a tumor-specific peptide sequences that single nucleotide variants and sequences from the of a human proteome database. was used to identify tumor and peptides to both germline and somatic A of germline and and luminal germline and variant peptides were by we only a to be both by MS and also at the mRNA The of variants by MS can be on peptide to provide MS including and peptide within peptides that to than proteomic are to to a J.P. The of peptide for protein PubMed Scopus Google Scholar). Therefore, we only tumor-specific with as variant peptides. that of variant peptides to than for and luminal which were from analysis. An of all variant peptides were to be of the peptide for both tumor The proportion of variants that are at the mRNA level, however, be and relies on variant from whole transcriptome the of peptides the for a for expression at the mRNA level germline germline 10% somatic somatic and A of variant peptides were identified across all sample process replicates, of which were by both and a of SNVs spectral of and global proteomics data using tumor-specific databases identified a of and luminal variant peptides both somatic and germline variant peptide at resulting from a with of the peptides with and peptide with An of variants identified were in and databases, only of identified protein variants by database. of the variants identified by proteomics analysis mRNA on RNA-Seq variant can be at by the with variant from RNA-Seq data to the of the transcriptome identification of variants from J. 2013; Full Text Full Text PDF PubMed Scopus Google Scholar). the variant peptides in the tumor, were to germline variants and only were to somatic variants of the somatic variants were identified by in the and in the In the luminal tumor, somatic and germline variant peptides were and than of the luminal somatic variants were at the mRNA The in somatic variant expression the and the luminal is differential translation or protein degradation in the of the variant peptides were only identified by only spectral match across all MS/MS or peptide spectral expression levels comparison of the proteomic variants with SNVs in of the The for analysis by that of the variants were at the DNA level in at of the human identified a of somatic variant peptides that than peptide spectral identified in at process RNA and for expression in the protein the variants were identified as tumor (somatic with including variants in the protein and the protein peptides In to determine the of sample process replicates that are to identify the of translated variants, we a analysis using global and MS/MS demonstrated that MS analysis likely in variant peptide identification and The of peptides identified in replicates to for and to for and the identification of variant peptides to the of peptides identified in label Despite this even process replicates with the peptide identified less than of the variant peptides and In to resulting from and novel expression have to tumor and A. J. B. The and of PubMed Scopus Google Scholar). by RNA-Seq analysis were used to identify to novel to only and novel to and in and luminal The of both the and was identified to determine the proteomic potential to novel splice junction (NSJ) peptides. of of novel and less than of novel peptides were from the analysis to their mapping to An of novel junction peptides were to be of the MS peptide on this we a of and and novel junction and and novel peptides to be by MS for luminal and databases for spectral we were to identify less than of the novel junction peptides for both and luminal all RNA-Seq junction as for protein without the of used this to create the with the MPS all proteomic with the that the proteomic data can be used to through is however, that identified by proteomics analysis will have therefore the of junction peptides and that novel with MS RNA sequencing luminal with all in the protein luminal of junction data or novel junction peptide to and luminal and of novel junction peptides identified by MS/MS this of the and of the luminal The of peptides identified to peptides was of the junction A of novel splice junction (NSJ) peptides were identified in by sample process the NSJ peptides identified in of a with a novel and demonstrated the of novel peptides of with and novel were identified in luminal of the NSJ peptides were identified in both and luminal of novel junction peptides were by only peptide spectral match by only process identified novel junction peptides we to be at peptide spectral than RNA and for protein expression in the In a NSJ peptide to be in both the and the luminal tumor, were observed for both and RNA-Seq for both In the sample is as both and the identification from analysis was to determine the sample of The peptide with the was for a novel in luminal novel junction peptides analysis for the variant peptides to the of novel peptides identified with process with the variant from all replicates were to the of junction peptides in both and analysis and The novel peptide identification was with peptides identified than that with the variant peptide with an of for and for tandem that the of novel and variant peptide identification is dependent on the of proteomic and in MS sensitivity will in comprehensive variant peptide understand this in and NSJ peptides with process we a for peptide identification for and label free MS/MS analysis a in peptide identification with process though the MS/MS that the of process replicates using and can identify a of novel peptides. and variants were in tumor-specific databases, peptides both a novel junction and were identified from analysis. also used to novel peptides from a translation of in and in A of luminal peptides and peptides were in the protein sequence search were using proteomics analysis. this in a with potential proteins in the protein sequence database. proteins were of the within a of protein expression the The of has in the a on the integration of and proteomics through peptide mapping and the of sequencing data to a comprehensive of the that MPS data to and identify proteins are essential for protein integration of sequencing and proteomic analysis will of biological in that comprehensive and data from the can provide a of biological multiple datasets can the with proteomics can be used to the protein of the genome N.E. Shen Z. He Y. Walley J.W. Cassidy C.J. Briggs S.P. Bafna V. An automated proteogenomic method utilizes mass spectrometry to reveal novel genes in Zea mays.Mol. Cell. Proteomics. 2013; 13: 157-167Abstract Full Text Full Text PDF PubMed Scopus (65) Google Scholar, Wang Y. D. X. D. A. P. A. A. W. and of transcriptome by parallel sequencing and shotgun Res. 2011; PubMed Scopus Google Scholar, He Y. Bafna V. from large data.J. Proteome Res. 2013; 13: PubMed Scopus Google Scholar, X. X. J. Y. The discovery of novel in genome on mass spectrometry 2011; PubMed Scopus Google DNA sequencing can protein variants in MS/MS J. Su Z. Ma Z.Q. Slebos R.J. Halvey P. Tabb D.L. Liebler D.C. Pao W. Zhang B. A bioinformatics workflow for variant peptide detection in shotgun proteomics.Mol. Cell. Proteomics. 2011; 10 (M110.006536)Abstract Full Text Full Text PDF Scopus (84) Google and RNA-Seq can MS proteomic coverage G.M. Shortreed M.R. Frey B.L. Smith L.M. Discovery and mass spectrometric analysis of novel splice-junction peptides using RNA-Seq.Mol. Cell. Proteomics. 2013; Full Text Full Text PDF PubMed Scopus Google Scholar, splice variants, a new of protein cancer in cancer and cancer with biology 2010; PubMed Google Scholar, The of mass proteomic data for of novel splice from RNA-Seq A 2010; 11: PubMed Scopus Google Scholar). study the current and of proteogenomic integration for cancer in the detection of novel protein the of current MS proteomic methods for variant and novel peptide identification in a sample that has both proteomic global MS/MS sample process and MPS analysis. the to than peptides that the expression of single nucleotide and novel is MS was to than of the of of the or for both variant and junction peptide analysis that process replicates to the of novel peptides the sensitivity of proteomics has the of the analysis is an variant peptides of cancer genes as or were not detected in this peptide and have large on and Y. J.M. Smith of the and of peptide PubMed Scopus Google peptides are for MS analysis are with protein coverage are using tandem MS In the of the analysis is of than the of protein even of peptides is is to a MS only of for a given and a in J. D. the of proteome analysis by and PubMed Scopus Google Scholar). the this is by sample in MS protein identification sensitivity be through in sample sensitivity and or identification we are to proteomic The novel junction peptides and identified in the proteomics data is and levels of discovery have in G.M. Shortreed M.R. Frey B.L. Smith L.M. Discovery and mass spectrometric analysis of novel splice-junction peptides using RNA-Seq.Mol. Cell. Proteomics. 2013; Full Text Full Text PDF PubMed Scopus Google Scholar). the that is a level of in the transcriptome, many in proteomics an essential technology in novel protein and splice also to the quality of the human with novel the in sequencing and the in novel junction peptides identified in the luminal and that peptides to be to a study also the current and in MPS analysis. A of the peptides novel were only by a or a RNA-Seq MS peptide many variants only quality from or not have RNA-Seq MS proteomics the method of for which variants will be proteins or protein and degradation and in the search for of The of translated somatic variants in the basal-like sample in comparison to the luminal tumor be in and degradation in to the of genes that are and of is essential in including SNVs and novel to the degree to which are translated and therefore peptide coverage is by current with only 10% coverage of variant peptide sequence even with multiple and process replicates, to in cancer using proteomics is variant peptide coverage for is clearly and the of and are The of peptide detection not be a sensitivity biological effects. the detection for novel the that many are not translated or of peptide expression for be to or protein The of proteogenomic integration methods to datasets and in peptide identification and MS/MS sensitivity will in the
Ruggles et al. (Wed,) studied this question.