Key points are not available for this paper at this time.
Shotgun proteomics data analysis usually relies on database search. However, commonly used protein sequence databases do not contain information on protein variants and thus prevent variant peptides and proteins from been identified. Including known coding variations into protein sequence databases could help alleviate this problem. Based on our recently published human Cancer Proteome Variation Database, we have created a protein sequence database that comprehensively annotates thousands of cancer-related coding variants collected in the Cancer Proteome Variation Database as well as noncancer-specific ones from the Single Nucleotide Polymorphism Database (dbSNP). Using this database, we then developed a data analysis workflow for variant peptide identification in shotgun proteomics. The high risk of false positive variant identifications was addressed by a modified false discovery rate estimation method. Analysis of colorectal cancer cell lines SW480, RKO, and HCT-116 revealed a total of 81 peptides that contain either noncancer-specific or cancer-related variations. Twenty-three out of 26 variants randomly selected from the 81 were confirmed by genomic sequencing. We further applied the workflow on data sets from three individual colorectal tumor specimens. A total of 204 distinct variant peptides were detected, and five carried known cancer-related mutations. Each individual showed a specific pattern of cancer-related mutations, suggesting potential use of this type of information for personalized medicine. Compatibility of the workflow has been tested with four popular database search engines including Sequest, Mascot, X!Tandem, and MyriMatch. In summary, we have developed a workflow that effectively uses existing genomic data to enable variant peptide detection in proteomics. Shotgun proteomics data analysis usually relies on database search. However, commonly used protein sequence databases do not contain information on protein variants and thus prevent variant peptides and proteins from been identified. Including known coding variations into protein sequence databases could help alleviate this problem. Based on our recently published human Cancer Proteome Variation Database, we have created a protein sequence database that comprehensively annotates thousands of cancer-related coding variants collected in the Cancer Proteome Variation Database as well as noncancer-specific ones from the Single Nucleotide Polymorphism Database (dbSNP). Using this database, we then developed a data analysis workflow for variant peptide identification in shotgun proteomics. The high risk of false positive variant identifications was addressed by a modified false discovery rate estimation method. Analysis of colorectal cancer cell lines SW480, RKO, and HCT-116 revealed a total of 81 peptides that contain either noncancer-specific or cancer-related variations. Twenty-three out of 26 variants randomly selected from the 81 were confirmed by genomic sequencing. We further applied the workflow on data sets from three individual colorectal tumor specimens. A total of 204 distinct variant peptides were detected, and five carried known cancer-related mutations. Each individual showed a specific pattern of cancer-related mutations, suggesting potential use of this type of information for personalized medicine. Compatibility of the workflow has been tested with four popular database search engines including Sequest, Mascot, X!Tandem, and MyriMatch. In summary, we have developed a workflow that effectively uses existing genomic data to enable variant peptide detection in proteomics. DNA sequence variation is associated with diseases and differential drug response. As a paradigmatic example, cancers are diseases of clonal proliferations caused by mutations in oncogenes and tumor suppressor genes (1Vogelstein B. Kinzler K.W. Cancer genes and the pathways they control.Nat. Med. 2004; 10: 789-799Crossref PubMed Scopus (3337) Google Scholar). After several decades of searching through traditional biology approaches, many mutant genes have been causally implicated in oncogenesis (2Futreal P.A. Coin L. Marshall M. Down T. Hubbard T. Wooster R. Rahman N. Stratton M.R. A census of human cancer genes.Nat. Rev. Cancer. 2004; 4: 177-183Crossref PubMed Scopus (2419) Google Scholar). Facilitated by the new genomic techniques such as SNP (single nucleotide polymorphism) arrays and deep-sequencing, the identification of cancer genes has made enormous progress over the past several years (3Wood L.D. Parsons D.W. Jones S. Lin J. Sjöblom T. Leary R.J. Shen D. Boca S.M. Barber T. Ptak J. Silliman N. Szabo S. Dezso Z. Ustyanksky V. Nikolskaya T. Nikolsky Y. Karchin R. Wilson P.A. Kaminker J.S. Zhang Z. Croshaw R. Willis J. Dawson D. Shipitsin M. Willson J.K. Sukumar S. Polyak K. Park B.H. Pethiyagoda C.L. Pant P.V. Ballinger D.G. Sparks A.B. Hartigan J. Smith D.R. Suh E. Papadopoulos N. Buckhaults P. Markowitz S.D. Parmigiani G. Kinzler K.W. Velculescu V.E. Vogelstein B. The genomic landscapes of human breast and colorectal cancers.Science. 2007; 318: 1108-1113Crossref PubMed Scopus (2577) Google Scholar, 4Weir B.A. Woo M.S. Getz G. Perner S. Ding L. Beroukhim R. Lin W.M. Province M.A. Kraja A. Johnson L.A. Shah K. Sato M. Thomas R.K. Barletta J.A. Borecki I.B. Broderick S. Chang A.C. Chiang D.Y. Chirieac L.R. Cho J. Fujii Y. Gazdar A.F. Giordano T. Greulich H. Hanna M. Johnson B.E. Kris M.G. Lash A. Lin L. Lindeman N. Mardis E.R. McPherson J.D. Minna J.D. Morgan M.B. Nadel M. Orringer M.B. Osborne J.R. Ozenberger B. Ramos A.H. Robinson J. Roth J.A. Rusch V. Sasaki H. Shepherd F. Sougnez C. Spitz M.R. Tsao M.S. Twomey D. Verhaak R.G. Weinstock G.M. Wheeler D.A. Winckler W. Yoshizawa A. Yu S. Zakowski M.F. Zhang Q. Beer D.G. Wistuba II Watson M.A. Garraway L.A. Ladanyi M. Travis W.D. Pao W. Rubin M.A. Gabriel S.B. Gibbs R.A. Varmus H.E. Wilson R.K. Lander E.S. Meyerson M. Characterizing the cancer genome in lung adenocarcinoma.Nature. 2007; 450: 893-898Crossref PubMed Scopus (929) Google Scholar, 5TCGA Comprehensive genomic characterization defines human glioblastoma genes and core pathways.Nature. 2008; 455: 1061-1068Crossref PubMed Scopus (5836) Google Scholar, 6Sjöblom T. Jones S. Wood L.D. Parsons D.W. Lin J. Barber T.D. Mandelker D. Leary R.J. Ptak J. Silliman N. Szabo S. Buckhaults P. Farrell C. Meeh P. Markowitz S.D. Willis J. Dawson D. Willson J.K. Gazdar A.F. Hartigan J. Wu L. Liu C. Parmigiani G. Park B.H. Bachman K.E. Papadopoulos N. Vogelstein B. Kinzler K.W. Velculescu V.E. The consensus coding sequences of human breast and colorectal cancers.Science. 2006; 314: 268-274Crossref PubMed Scopus (2857) Google Scholar, 7Greenman C. Stephens P. Smith R. Dalgliesh G.L. Hunter C. Bignell G. Davies H. Teague J. Butler A. Stevens C. Edkins S. O'Meara S. Vastrik I. Schmidt E.E. Avis T. Barthorpe S. Bhamra G. Buck G. Choudhury B. Clements J. Cole J. Dicks E. Forbes S. Gray K. Halliday K. Harrison R. Hills K. J. A. Jones D. A. T. J. K. D. Shepherd R. A. C. J. T. S. S. A. P. F. L. A. M.F. P. E. G. Wooster R. P.A. Stratton M.R. of in human cancer 2007; PubMed Scopus Google Scholar). The genomic of cancer are through proteins and and proteins the genomic in cancer have the potential to discovery and has to As a have into the past shotgun proteomics has as a for the identification of proteins in C.L. Zhang Y. Zhang Y. M. A by protein 2006; PubMed Scopus Google Scholar, T. B. A. C. P. A. M.S. Q. J. B. A. of and protein in and 2006; PubMed Scopus Google Scholar). to tumor potential in mutant proteins in human However, shotgun proteomics data analysis usually relies on database search and commonly protein sequence databases do not contain protein variation the of shotgun proteomics to the detection of protein sequence variants a have made on the identification of variant peptides on the search of sequence A modified of search of human variants through variations and then a database that sequences to peptides C.L. J.K. J.R. identification of sequence variations in proteins by PubMed Scopus Google Scholar). Roth Forbes Robinson and characterization of coding and in human proteins by 4: PubMed Scopus Google developed a human protein database for the by of protein in a search. the search in J.S. searching of PubMed Scopus Google and the search in R. proteins with 2004; PubMed Scopus Google of that from nucleotide in of the search is to of for the variant identifications and the J.S. searching of PubMed Scopus Google Scholar). to the search of protein variants is to from known coding A SNP was by in were protein databases and a SNP database created from peptides from the for database J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). the protein sequence database through the sequences with from sequence and peptides S. J. B. Zhang Y. J.S. M. A database for 2007; 4: PubMed Scopus Google Scholar). a was created for human mutant sequences on the search of shotgun proteomics data H. Park J. Ding G. Y. the for human sequences from PubMed Scopus Google Scholar). human mutant sequences from the in A. A.F. J.S. in a of human genes and PubMed Scopus Google Database T. M. K. The PubMed Scopus Google and database B. A. R. A. E. K. C. I. S. M. The protein and in PubMed Scopus Google Scholar). the of shotgun proteomics to the identification of protein variants in human cancers has not been addressed mutations, are not in existing database a of genome variation to by has been for to the of cancers M. B. R. A. M. A. M. E. A. N. R. a for sequence and for variation in 2004; PubMed Google Scholar). However, cancer mutations are collected in the of in Cancer S. Dawson E. Forbes S. Clements J. R. A. A. Teague J. P.A. Stratton M.R. Wooster R. The of in database and J. Cancer. 2004; PubMed Scopus Google and cancer specific databases M. A. Teague J. Forbes S. J.K. A. M. R.G. P. databases as for and of for data and PubMed Scopus Google As a mutations have been from we developed a human Cancer Proteome Variation database used human Cancer Proteome Variation nucleotide J. Zhang B. a human cancer variation PubMed Scopus Google that comprehensively variation data from a of cancer specific variation data including B. L. L. B. A. and in R. PubMed Scopus Google Scholar, C. R. A. The human proteomics PubMed Scopus Google and on cancer genes and cancer T. Jones S. Wood L.D. Parsons D.W. Lin J. Barber T.D. Mandelker D. Leary R.J. Ptak J. Silliman N. Szabo S. Buckhaults P. Farrell C. Meeh P. Markowitz S.D. Willis J. Dawson D. Willson J.K. Gazdar A.F. Hartigan J. Wu L. Liu C. Parmigiani G. Park B.H. Bachman K.E. Papadopoulos N. Vogelstein B. Kinzler K.W. Velculescu V.E. The consensus coding sequences of human breast and colorectal cancers.Science. 2006; 314: 268-274Crossref PubMed Scopus (2857) Google Scholar, 7Greenman C. Stephens P. Smith R. Dalgliesh G.L. Hunter C. Bignell G. Davies H. Teague J. Butler A. Stevens C. Edkins S. O'Meara S. Vastrik I. Schmidt E.E. Avis T. Barthorpe S. Bhamra G. Buck G. Choudhury B. Clements J. Cole J. Dicks E. Forbes S. Gray K. Halliday K. Harrison R. Hills K. J. A. Jones D. A. T. J. K. D. Shepherd R. A. C. J. T. S. S. A. P. F. L. A. M.F. P. E. G. Wooster R. P.A. Stratton M.R. of in human cancer 2007; PubMed Scopus Google Scholar). coding variations in are in variation to a protein sequence database that protein variant detection in shotgun proteomics analysis of human cancer the human Cancer Proteome Variation Database nucleotide protein variants to known coding and mutations could effectively the search as with the of this the of in a protein sequence database, in the risk of false positive to this J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). In the by J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google a peptide is as SNP the search for the is the for The of was on to the false and false J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). was in this of the by and and by variations to sequence databases of variation information in the database, of the database with search and of that variant and In this we workflow to the we created a protein sequence database on the we developed a workflow for and variant peptides from shotgun proteomics We used data sets from colorectal cancer cell lines and human to our variants were through genomic sequencing. we tested the of the workflow with popular search engines including peptide identification by Proteome 2007; PubMed Scopus Google J.K. J.R. to data of peptides with sequences in a protein PubMed Scopus Google J.S. protein identification by searching sequence databases PubMed Scopus Google and R. proteins with 2004; PubMed Scopus Google Scholar). A was developed to on the from search we our workflow the The human proteomics from colorectal cell lines SW480, and and three colorectal tumor were in the The cell lines were from and and of of or from that been made of were in and and with was in HCT-116 and were in were to the was were in and collected in were for and was were cell could carried were from the cell were and through the analysis tumor were from the colorectal cancer that from the We three on of the and confirmed for the of tumor by a A total of for of the was and collected into have been in R.J. M.A. M. of of peptides for Proteome 2008; PubMed Scopus Google Scholar, M. R.J. of protein from and in shotgun PubMed Scopus Google Scholar). In summary, proteins from cell or were with and with The peptides were on that were into cell or human Each of was by a on a by analysis on data in the were to the the in the D. M. R. D. P. for proteomics 2008; PubMed Scopus Google Scholar). variation data were from the database on in proteins and cancer-related variations in proteins J. Zhang B. a human cancer variation PubMed Scopus Google Scholar). A protein database was from We tested our workflow four popular database search including peptide identification by Proteome 2007; PubMed Scopus Google J.K. J.R. to data of peptides with sequences in a protein PubMed Scopus Google J.S. protein identification by searching sequence databases PubMed Scopus Google and R. proteins with 2004; PubMed Scopus Google was used as the search in this were to and were to A of to was were to of identifications that to three or peptide sequences with were was and was The for search engines are in DNA from cell lines RKO, SW480, and HCT-116 was a After identification of variant peptides by shotgun the the protein sequences were a The were by of and a of A of the used for the is in and were were by and then on sequence were in and As in our workflow for and variant peptides on shotgun proteomics data three database peptide and The protein sequence database was created on the protein database and the database J. Zhang B. a human cancer variation PubMed Scopus Google Scholar). variations and and were in the After the in cancer-related variation in was with the sequence the peptide and the peptides was as in the with were they in shotgun proteomics. the peptides for the identification of peptides with S. J. B. Zhang Y. J.S. M. A database for 2007; 4: PubMed Scopus Google Scholar). database the as sequence variants to the protein was in the of S. J. B. Zhang Y. J.S. M. A database for 2007; 4: PubMed Scopus Google Scholar). We to peptides as variation information in the sequence protein the and of the peptide in the and the of the variation in database or new peptide in of in the peptide database protein sequence database the protein database and peptide with variations from peptide carried cancer-related variations. We this protein sequence database sequences were as sequences for false discovery rate estimation search for in protein identifications by 2007; 4: PubMed Scopus Google Scholar). After of the database, shotgun proteomics data from a cancer the database a database search The is the of the peptide has been that a risk of false could associated with variant peptide identifications as with that for ones J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). In to this we the with and used the estimation search for in protein identifications by 2007; 4: PubMed Scopus Google with to variant with were into a and a variant and the of were The for the variant showed a the were in data from cancer cell lines not that the of peptides were the the variant have a risk of false positive genomic analysis further confirmed this As in the variant peptides randomly with were confirmed with genomic sequencing. a of the rate was of for the variant peptides in and to estimation and the estimation in this with the are with and or genomic in a new sequences are as and for a specific of the sequences are the of sequences variant sequences in the database is to that variant sequences are to have a of in a specific However, the estimation is for the selected from variant the estimation to a false rate for peptide identifications and a false positive rate for variant peptide not a for identifications variant sequences a of the we variant peptide the for the could the this we the for and variant peptides variant peptides and were for the estimation of variant peptide this were for the variant peptides for the ones in of the and the risk of high false for the variant was to the However, in a search was for the variant peptides that for the for variant we that the search to a the a we not in the of the of in the variant peptides and variant In the search for the total of false that a specific by the of selected the of with the the of total is for a of the of false As a on the genomic rate was a estimation of the total of false we to information on from variant and sequences and for variant identifications on the and are the of and variant the and are the of and variant the The of the of variant sequences in the In this the of false in variant identifications is by the total of false and the of variant sequences in the is that is the from sequences and variant sequences for a estimation of the of false for variant peptide identifications the is on data that are to The showed that the new estimation could the of variant peptide identifications for the variations showed that the new estimation and As in variant peptide identification on the new a rate of of as with the rate of of on or the was not confirmed the genomic this through the of the the of the new estimation for variant peptide identifications was in our workflow In the of the and variant peptides are on the estimation and is the we database search and peptide identification for three data sets from colorectal cancer cell lines RKO, and SW480, was used as the search and the was to for and variant and peptides were in SW480, RKO, and HCT-116 were to and protein B. through analysis and Proteome 2007; PubMed Scopus Google Scholar, S. S.M. B. protein with high peptide identification Proteome PubMed Scopus Google Scholar). The of variant peptides was and for SW480, RKO, and and to and of peptides in cell We randomly selected and variant peptides from the and HCT-116 data sets for genomic and the rate were of and of the genomic for SW480, the rate for three cell lines was A of the variant peptides and associated information in In the HCT-116 data we a variation in was of the genes as a of tumor in The variation is not a known in the HCT-116 cell has been in of colorectal in a C. D. M. E. A. S. R. K. P. P. I. M.R. M. K. P. R. C. S. A. P. H. L.A. R. of mutations in colorectal to and 2004; PubMed Scopus Google Scholar). In to the cancer cell we applied the on three data sets from colorectal tumor specimens. A total of peptides were in data were to protein The of distinct variant peptides in was 204 and to of As in and variant peptides were in of the three peptides carried known cancer-related mutations, and of were in The was in colorectal cancer in this in are the commonly mutations in genes with of human cancers mutations in this tumor suppressor T. in human the 2007; PubMed Scopus Google Scholar). The of mutations of a protein to of cell to cell and the of in such mutations. The has been in the cell mutant in cell in and in and to G. E. S. C. G. A. of of tumor of human cancer cell lines through of mutant 2006; PubMed Scopus Google Scholar, W. Liu G. A. of of a of mutant is for 2008; PubMed Scopus Google Scholar). The was in not in colorectal cancer this has been in lung cancer Bhamra G. S. Dawson E. C. Clements J. A. Teague P.A. Stratton M.R. The of in Cancer 2008; PubMed Scopus Google Scholar). A of mutations were in several cell lines from of the and and has been as a drug for cancer F. Y. L. P. K. S. D. B. Smith R. G. A. K. J. B. E. a of the is in human tumor cell Google Scholar). As a of is a of and has been to for the of The is a of for of PubMed Scopus Google Scholar). our workflow and popular proteomics search we tested the with Sequest, Mascot, as well as MyriMatch. The of search engines are in the or into the The variation information for variant peptide is the is specific to this we created a that used to identification for variant and and peptide information for information in a protein and The is in and from our on the data and distinct variant peptides Sequest, Mascot, and X!Tandem, variant peptides were by or search engines is not to from search and from search engines has been as a to peptide identification J.A. Hubbard in by analysis of false discovery for search PubMed Scopus Google Scholar, M. by from search Proteome 2008; PubMed Scopus Google Scholar, W. J.A. S.D. the and of peptide identification in by search 10: PubMed Scopus Google Scholar). The to use our with search engines to this type of on search for that from nucleotide in the search in J.S. searching of PubMed Scopus Google and the search in R. proteins with 2004; PubMed Scopus Google the detection of variant peptides existing information on genomic sequence variations. our we analysis on the data and the protein We the a for the peptide variant peptide we the from they have of the for and they have in of the to the the search in peptides and variant the search in peptides and variant we the variant peptides by search the search engines and and of the identifications were to and X!Tandem, In the the search engines our was with identifications for search The of search from the search engines a of high false positive many variant peptides our we identifications from the search suggesting a of into the of the we variant identifications the variant peptides by in the cell and confirmed by genomic the of variant peptide for search three out of the peptides were In with our out of the were by and a made on the of do not for a of the We have created a protein database and a workflow for the identification of and variant peptides on the A estimation was in the workflow to high of the variant of variant peptides were from three colorectal cancer cell lines and three tumor used in this of the variants were from the database and are to are associated with cancer known cancer-related mutations have been including associated with cell and drug A on the use of databases for shotgun proteomics data searching is the high risk of false positive identifications J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). In this we this risk by the search of and variant peptide identifications and a modified estimation to this existing of a for variant peptide identifications J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). is that in our were carried out for variant and peptide the database search was the to S. J. B. Zhang Y. J.S. M. A database for 2007; 4: PubMed Scopus Google Scholar). In J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google were for the and variant a variant database is a to a variant peptide of the of the from the was used to of the of the peptide variants our workflow and of were false by the proteomics identifications and genomic by of peptide J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). example, the in the data of as this was not confirmed by genomic sequencing. and are on peptides J.R. E. and of coding from analysis of shotgun proteomics Proteome 2007; PubMed Scopus Google Scholar). a has from a sequence variation or for example, the in could by the of the this was confirmed the genomic As out by S. J. B. Zhang Y. J.S. M. A database for 2007; 4: PubMed Scopus Google searching protein databases a new to proteomics In cancer detection of mutant peptides and proteins of individual by proteomics techniques have on the of personalized medicine. five known cancer-related mutations were in the tumor from three colorectal cancer showed a specific pattern and information a that could personalized cancer As with protein variants to from known coding and mutations could effectively the search and thus to However, this a of on known genomic sequence variations. The of known cancer-related mutations in this was this by the of is the database mutations in the database in human proteins have cancer-related mutations cancer genes have mutations in many in protein such as and a cancer-related mutations have been in by the of cancer of of cancer genome such as the Cancer of the The Cancer of the Cancer and the our on mutations in human cancers S. Dawson E. Forbes S. Clements J. R. A. A. Teague J. P.A. Stratton M.R. Wooster R. The of in database and J. Cancer. 2004; PubMed Scopus Google Scholar, Bhamra G. S. Dawson E. C. Clements J. A. Teague P.A. Stratton M.R. The of in Cancer 2008; PubMed Scopus Google Scholar). We from into and to the of our analysis many variant in this on not for a of over our a sequence has been as for variant peptide identification S. R.J. identification through sequence Proteome PubMed Scopus Google Scholar). to a of in to distinct variations and and were in protein sequence variations such as variants and are in cancer and have been by shotgun proteomics R. characterization of variant proteins in human breast PubMed Scopus Google Scholar, J. J. A for protein analysis and 2006; PubMed Scopus Google Scholar). is to our database and workflow for the of existing on variations. In summary, we have developed a workflow for variant peptide detection in shotgun proteomics The workflow a variation detection and of peptide Compatibility of the workflow with popular database search engines has been of the identifications has been confirmed by genomic sequencing. this workflow on human cancer proteomics into cancer and potential personalized We to and C. for in the of the with
Li et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: