Key points are not available for this paper at this time.
The original PRIDE Converter tool greatly simplified the process of submitting mass spectrometry (MS)-based proteomics data to the PRIDE database. However, after much user feedback, it was noted that the tool had some limitations and could not handle several user requirements that were now becoming commonplace. This prompted us to design and implement a whole new suite of tools that would build on the successes of the original PRIDE Converter and allow users to generate submission-ready, well-annotated PRIDE XML files. The PRIDE Converter 2 tool suite allows users to convert search result files into PRIDE XML (the format needed for performing submissions to the PRIDE database), generate mzTab skeleton files that can be used as a basis to submit quantitative and gel-based MS data, and post-process PRIDE XML files by filtering out contaminants and empty spectra, or by merging several PRIDE XML files together. All the tools have both a graphical user interface that provides a dialog-based, user-friendly way to convert and prepare files for submission, as well as a command-line interface that can be used to integrate the tools into existing or novel pipelines, for batch processing and power users. The PRIDE Converter 2 tool suite will thus become a cornerstone in the submission process to PRIDE and, by extension, to the ProteomeXchange consortium of MS-proteomics data repositories. The original PRIDE Converter tool greatly simplified the process of submitting mass spectrometry (MS)-based proteomics data to the PRIDE database. However, after much user feedback, it was noted that the tool had some limitations and could not handle several user requirements that were now becoming commonplace. This prompted us to design and implement a whole new suite of tools that would build on the successes of the original PRIDE Converter and allow users to generate submission-ready, well-annotated PRIDE XML files. The PRIDE Converter 2 tool suite allows users to convert search result files into PRIDE XML (the format needed for performing submissions to the PRIDE database), generate mzTab skeleton files that can be used as a basis to submit quantitative and gel-based MS data, and post-process PRIDE XML files by filtering out contaminants and empty spectra, or by merging several PRIDE XML files together. All the tools have both a graphical user interface that provides a dialog-based, user-friendly way to convert and prepare files for submission, as well as a command-line interface that can be used to integrate the tools into existing or novel pipelines, for batch processing and power users. The PRIDE Converter 2 tool suite will thus become a cornerstone in the submission process to PRIDE and, by extension, to the ProteomeXchange consortium of MS-proteomics data repositories. The sharing of biological data in the public domain is generally considered to be good scientific practice. This concept of data sharing has gained substantial traction in the field of MS-based proteomics, in which the PRIDE 1The abbreviations used are:APIApplication Programming InterfaceBBSRCBiotechnology and Biological Science Research CouncilCLICommand-Line InterfaceCVcontrolled vocabularyDAOdata access objectEBIEuropean Bioinformatics InstituteGUIgraphical user interfaceJARjava archiveLIMSLaboratory Information Management SystemMIAPEMinimum Information About a Proteomics ExperimentNIHNational Institutes of HealthOLSOntology Lookup ServicePMFpeptide mass fingerprintingPRIDEPRoteomics IDEntifications databasePSIProteomics Standards InitiativePTMpost-translational modificationPXProteomeXchangeUniProtKBUniProt KnowledgeBaseXMLeXtensible Markup Language. 1The abbreviations used are:APIApplication Programming InterfaceBBSRCBiotechnology and Biological Science Research CouncilCLICommand-Line InterfaceCVcontrolled vocabularyDAOdata access objectEBIEuropean Bioinformatics InstituteGUIgraphical user interfaceJARjava archiveLIMSLaboratory Information Management SystemMIAPEMinimum Information About a Proteomics ExperimentNIHNational Institutes of HealthOLSOntology Lookup ServicePMFpeptide mass fingerprintingPRIDEPRoteomics IDEntifications databasePSIProteomics Standards InitiativePTMpost-translational modificationPXProteomeXchangeUniProtKBUniProt KnowledgeBaseXMLeXtensible Markup Language. (PRoteomics IDEntifications) database (http://www.ebi.ac.uk/pride) at the European Bioinformatics Institute (EBI, Cambridge, UK) is one of the most prominent public data repositories (1Vizcaíno J.A. Côté R. Reisinger F. Barsnes H. Foster J.M. Rameseder J. Hermjakob H. Martens L. The Proteomics Identifications database: 2010 update.Nucleic Acids Res. 2010; 38: D736-742Crossref PubMed Scopus (211) Google Scholar). PRIDE stores MS and MS/MS spectra, the derived peptide and protein identifications and expression values if available (the processed experimental results), and any associated metadata. It is important to highlight that data stored in PRIDE is not reprocessed after submission. PRIDE, in its current form, represents the submitter's view of the data. PRIDE is also a founding member of the ProteomeXchange (PX) consortium (http://www.proteomexchange.org) (2Hermjakob H. Apweiler R. The Proteomics Identifications Database (PRIDE) and the ProteomExchange Consortium: making proteomics data accessible.Expert Rev. Proteomics. 2006; 3: 1-3Crossref PubMed Scopus (71) Google Scholar). The PX members, led by PRIDE and PeptideAtlas (3Deutsch E.W. Lam H. Aebersold R. PeptideAtlas: a resource for target selection for emerging targeted proteomics workflows.EMBO Rep. 2008; 9: 429-434Crossref PubMed Scopus (444) Google Scholar), are currently working toward the implementation of a system that enables the automated and standardized sharing of MS-based proteomics data between the main proteomics repositories. In this framework, PRIDE is the initial submission point for tandem MS data. Currently, the first pilot PX submissions (containing raw data and processed results) have already been carried out (http://proteomecentral.proteomexchange.org) and the system is now starting to accept regular submissions. At present, submissions to PRIDE are performed using a publicly available XML data format called PRIDE XML, which is built around the mzData data standard format (4Orchard S. Montechi-Palazzi L. Deutsch E.W. Binz P.A. Jones A.R. Paton N. Pizarro A. Creasy D.M. Wojcik J. Hermjakob H. Five years of progress in the Standardization of Proteomics Data 4th Annual Spring Workshop of the HUPO-Proteomics Standards Initiative April 23–25, 2007 Ecole Nationale Superieure (ENS), Lyon, France.Proteomics. 2007; 7: 3436-3440Crossref PubMed Scopus (46) Google Scholar). Application Programming Interface Biotechnology and Biological Science Research Council Command-Line Interface controlled vocabulary data access object European Bioinformatics Institute graphical user interface java archive Laboratory Information Management System Minimum Information About a Proteomics Experiment National Institutes of Health Ontology Lookup Service peptide mass fingerprinting PRoteomics IDEntifications database Proteomics Standards Initiative post-translational modification ProteomeXchange UniProt KnowledgeBase eXtensible Markup Language. Application Programming Interface Biotechnology and Biological Science Research Council Command-Line Interface controlled vocabulary data access object European Bioinformatics Institute graphical user interface java archive Laboratory Information Management System Minimum Information About a Proteomics Experiment National Institutes of Health Ontology Lookup Service peptide mass fingerprinting PRoteomics IDEntifications database Proteomics Standards Initiative post-translational modification ProteomeXchange UniProt KnowledgeBase eXtensible Markup Language. Several scientific journals (e.g. Molecular and Cellular Proteomics, Proteomics, and Nature Publishing Group journals) are supporting a gradual move toward mandating public deposition of MS data to support the publication of related manuscripts. In parallel, several funding agencies (such as The Wellcome Trust, NIH, and BBSRC) are also enforcing the public availability of experimental data in the context of their funded projects. Despite these efforts, the field of MS proteomics is still lagging behind other more mature “omics” disciplines in terms of public data availability (5Credit where credit is overdue.Nat. Biotechnol. 2009; 27 (No authors listed): 579Crossref PubMed Scopus (27) Google Scholar). In practical terms, a major contribution to this public data-sharing policy trend is provided by the availability of reliable and user-friendly submission tools. Such tools must be able to capture properly the experimental data and any supporting technical and biological metadata. In addition, to encourage MS data deposition the submission process has to be as easy as possible. This was the philosophy that drove the development of the original PRIDE Converter (6Barsnes H. Vizcaíno J.A. Eidhammer I. Martens L. PRIDE Converter: making proteomics data-sharing easy.Nat. Biotechnol. 2009; 27: 598-599Crossref PubMed Scopus (149) Google Scholar) (http://pride-converter.googlecode.com), an open source and platform-independent software tool for the submission of proteomics data to PRIDE. PRIDE Converter can convert input data from a large variety of popular MS proteomics formats into PRIDE XML, guiding the user through the process by a graphical user interface (GUI). As a result, PRIDE Converter made the submission of MS data a much easier and more straightforward process, especially for researchers without bioinformatics support. PRIDE Converter has definitely been a key factor in the huge growth in data content in PRIDE since 2008 (7Csordas A. Ovelleiro D. Wang R. Foster J.M. Ríos D. Vizcaíno J.A. Hermjakob H. PRIDE: quality control in a proteomics data repository.Database. 2012; 2012: bas004Crossref PubMed Scopus (35) Google Scholar) and has become the de facto submission tool to PRIDE for most researchers. PRIDE Converter has been regularly updated and more than 30 different releases have been made publicly available. However, after receiving extensive feedback from users, it became apparent that the original PRIDE Converter had some limitations mainly in terms of software architecture, memory requirements, difficulties to extend the supported formats, and a lack of functionality for performing batch conversions (a frequent request). In addition, new use cases needed to be supported, such as support for quantitative information and the ability to easily post-process the large XML files generated during the conversion process. To overcome these limitations, we decided to design a new submission tool from the ground up, which would be suitable to the evolving needs of our submitters. In this manuscript we describe the PRIDE Converter 2 framework, including all of its new features and supported use cases. We are that to PRIDE and to the PX consortium will from the availability of this new submission The PRIDE Converter 2 tool suite is in and all the source is available It is as open source the The development of PRIDE Converter 2 had several a of tools to tool be through a command-line interface for into PRIDE XML and tool be through a to a user-friendly tool has to be as as in its use of to a memory as input formats as by existing Application Programming and the as as in the tool suite by where on the original PRIDE Converter tool and support new use cases by our users. To these the PRIDE Converter 2 tool suite from the in which that are to a can be to resource As the PRIDE Converter 2 tool suite of different PRIDE Converter PRIDE mzTab PRIDE XML and PRIDE XML All of these tools are a for can the by on the or by it without from the are the is batch processing and, as a key into existing and built and are provided as a PRIDE Converter 2 for users, and the technical implementation can be in in the PRIDE Converter 2 tool Converter search files into well-annotated PRIDE XML files for mzTab skeleton mzTab files where the user can quantitative XML several PRIDE XML files in and peptide XML PRIDE XML files to to empty the protein in a new The original PRIDE Converter had some and practical the behind the development of the PRIDE Converter 2 tool suite was the to not overcome the of the original tool also functionality that had been by our users. As the PRIDE Converter 2 tool suite has conversion support for a of new data formats and other formats will support for existing formats has also been it is now to submit peptide mass data generated by the of quantitative data to PRIDE XML files has been greatly by support for mzTab formats in PRIDE Converter in PRIDE Converter and and F. R. F. Ríos D. Hermjakob H. Vizcaíno J. Jones A.R. interface to the standard for peptide and protein 2012; PubMed Scopus (27) Google and Barsnes H. Martens L. A. an to and MS/MS search 2010; PubMed Scopus Google and and and and and N. Barsnes H. A. Martens L. an open source to and Res. PubMed Scopus Google Reisinger F. Martens L. an for the standard for MS 2010; PubMed Scopus Google J. Reisinger F. Hermjakob H. Vizcaíno J.A. to process and and mass spectrometry data 2012; PubMed Scopus (27) Google J. Reisinger F. Hermjakob H. Vizcaíno J.A. to process and and mass spectrometry data 2012; PubMed Scopus (27) Google J. Reisinger F. Hermjakob H. Vizcaíno J.A. to process and and mass spectrometry data 2012; PubMed Scopus (27) Google J. Reisinger F. Hermjakob H. Vizcaíno J.A. to process and and mass spectrometry data 2012; PubMed Scopus (27) Google J. Reisinger F. Hermjakob H. Vizcaíno J.A. to process and and mass spectrometry data 2012; PubMed Scopus (27) Google Scholar) in a new The mzTab format is to be a standard for MS-based proteomics data, by the Proteomics Standards Initiative to be easy to it the information to the of a proteomics can generate skeleton mzTab files using the PRIDE mzTab and use the mzTab files as a basis to quantitative information as of the conversion process in PRIDE Converter and information can also be to the mzTab making the capture of information much more straightforward can now also their original search in format This is to data for protein and it easier to the all protein a process that is performed as a of in the PRIDE database to search J.A. Côté R. Reisinger F. Foster J.M. Rameseder J. Hermjakob H. Martens L. to the Proteomics Identifications Database proteomics data 2009; 9: PubMed Scopus Google Scholar). user by the PRIDE Converter 2 tool suite is the ability to post-process the generated PRIDE XML files. users can now use the PRIDE XML tool to contaminants and empty to submission. in the of gel-based proteomics in which one MS the original PRIDE Converter tool would generate one PRIDE XML This that a could several if not of PRIDE The PRIDE XML can now an large of PRIDE XML files into a the between and their This that users will be able to a PRIDE to to their experimental data. users will be by the user-friendly This interface has the of a and provides feedback on the and that are at of the conversion process. the user will be on to the the any (such as empty or data in the user is and the conversion process is the is by the a is out the user can be that the information is The process can also generate which would not the process still be to generate PRIDE XML files. The interface is mainly toward power users have the to batch conversions already have to all of the to and PRIDE XML files. using the the PRIDE Converter 2 tool must be in and to The needs to be first and will the files from an MS or without peptide and protein and will generate an from of files can be provided in the to the the protein search database used in the proteomics and mzTab files quantitative data and the has been properly the PRIDE Converter 2 tool is in to generate a PRIDE XML PRIDE Converter 2 currently the formats in formats can easily be supported by the This interface provides to access and information on spectra, and post-translational from the source files. The to as much information as from the source files to a starting point for the process. to is available at and in the To to for the formats, PRIDE Converter 2 use of existing where available this was not new were The generated by the will all of the protein and peptide and will as the basis for all controlled vocabulary that will its way into the PRIDE XML not on and software search database protein and the will also any quantitative and gel-based PRIDE Converter 2 will to protein from into the PRIDE format for that source as of the process the protein will be to which is the protein format in PRIDE for The XML is and and a has been provided to generate making it easy to integrate this functionality into an existing proteomics as a first in data to PRIDE the has been generated and or by using the PRIDE Converter 2 the of PRIDE Converter is to generate a PRIDE XML in PRIDE Converter 2 the user through a process to convert their search files into well-annotated PRIDE XML files format selection is the first that the user must can have one or several that can be through the are provided and the are by users have the to all available and the for these will be stored in the such that the user can the conversion process was process of search result files into well-annotated PRIDE XML files. The of the different in the conversion process are to an the to a submission. The not are that the of these on the of the input files. The other are related to selection and are of the of the files. of was used to the To users the format of their search the to convert and any if The process and on to and software processing The users are to or the and any The files are and the users can the process or to PRIDE XML This is where filtering can also be the conversion process has the users are to their PRIDE XML files the PRIDE tool and submit to PRIDE and to the ProteomeXchange The allows users to convert source files to the This can a of on software and search will be related source files. PRIDE Converter 2 also provides the to used such as and as which can be in of is provided the tool suite that users can to their allows the user to and using that the most used values in PRIDE. still have the to use the Ontology Lookup Service R. Reisinger F. Martens L. Barsnes H. J.A. Hermjakob H. The Ontology Lookup and Acids Res. 2010; 38: PubMed Scopus Google H. Eidhammer I. Martens L. to the Ontology Lookup 2010; PubMed Scopus Google Scholar) to terms that are not already the user has quantitative data an mzTab the will be used to for the features the of most in proteomics to the controlled vocabulary terms in the L. R. Binz P.A. J. Creasy D. J. The standard for of protein modification Biotechnol. 2008; PubMed Scopus Google Scholar). has been a in the and a source of in PRIDE data (7Csordas A. Ovelleiro D. Wang R. Foster J.M. Ríos D. Vizcaíno J.A. Hermjakob H. PRIDE: quality control in a proteomics data repository.Database. 2012; 2012: bas004Crossref PubMed Scopus (35) Google Scholar), as most search in different using PRIDE Converter 2 to a standardized on a of the most and the mass as by the search a can be to a mass a mass the is to the In cases where can be to a mass a of PRIDE Converter 2 will to a to a is at the it will be the will the that have been In the that are still at the is The will the that have been the mass by the in that have not been will be in The user must the to the using the if or by the for the for more is also performed in the are in the and not the It is to the user to the to the all of the are the a to the all files. The is the of the PRIDE XML files. the graphical conversion process can be at this as all of the files are now and and the conversion process can be and using the This is generally practical for users to convert a large of files and have access to a where the conversions can be In most a memory and will be more than The of the users to the generated PRIDE XML files using the PRIDE tool R. A. Ríos D. Ovelleiro D. Foster J.M. Côté J. A. Reisinger F. Hermjakob H. Martens L. Vizcaíno J.A. PRIDE a tool to and MS proteomics Biotechnol. 2012; PubMed Scopus Google Scholar) and to submit their data the PX to to of the for a user for all the tools in the PRIDE Converter 2 tool The PRIDE mzTab will generate skeleton mzTab files on the MS source files used by PRIDE Converter The user has the as in PRIDE Converter 2 and the are also stored in the mzTab This is important as the mzTab files and the files to be generated the as the conversion to and mzTab files are used as of the PRIDE Converter 2 the of the mzTab files are and if not the for the an is to the user and conversion is the are The PRIDE mzTab has several to handle and quantitative the quantitative it is for the PRIDE mzTab to to describe the used in the and to the for the that the users will be able to to the the is it is to to a on a on This information can also be from the if it is All of this information will be stored in the and will its way into the PRIDE XML It is a in MS-based proteomics that several files are from a of this would be a gel-based MS in which a MS and associated It has already been that PRIDE Converter 2 is able to these input files and convert in one a of for all of the source files. The PRIDE XML is the in such a where all the files are into a XML for submission. the PRIDE XML it is to generate one PRIDE XML which is a more than one PRIDE XML The PRIDE XML is to post-process the PRIDE XML files generated by PRIDE Converter 2 and at the of protein identifications and To the PRIDE XML can empty of or protein identifications that than a of to for The PRIDE XML can also a of protein identifications and use this as a to identifications from the XML The protein is one of the major in proteomics Aebersold R. of the protein Proteomics. PubMed Scopus Google Scholar). the PRIDE XML format not support properly the of to and these into by PRIDE Converter 2 all of peptide to protein making that data is This the of we have a into the PRIDE XML which can a of generated using an protein and all from the generated PRIDE XML this is not an we that it is a between the limitations of PRIDE XML and the of protein to the of the for a more on the protein The PRIDE Converter 2 a the original PRIDE Converter submission The behind the original tool has been the software must be as user-friendly as for without much bioinformatics support. that original the now use cases that were from the original tool were much in users. As a result, PRIDE Converter 2 can now be used by to batch it can be into to submissions to PRIDE, and it and quantitative data submissions. As of PRIDE Converter 2 has already been used to generate more than PRIDE XML different input We that the original PRIDE Converter will be in the The software architecture, and availability of the source allow any to support for a new format by a suitable implementation of the In this has already in the as the supporting data from files was of the PRIDE using the existing N. Barsnes H. A. Martens L. an open source to and Res. PubMed Scopus Google Scholar), in the context of the We that the new features we have to the PRIDE Converter 2 tool such as into batch conversion of and of new formats, will proteomics to their submission into PRIDE or integrate to PRIDE XML files in other tools. the we would encourage other have data formats not currently supported to conversion for PRIDE Converter We that the requirements for data availability by scientific journals and funding agencies can be in a much more and user-friendly way by this new As well as support for new formats, we are that the open of PRIDE Converter 2 will encourage to a for the generated PRIDE XML files. for the L. S. Reisinger F. Jones A.R. Martens L. Hermjakob H. The a to of proteomics 2009; 9: PubMed Scopus Google Scholar) is already into the PRIDE Converter 2 and we have to that would user such as requirements and the PRIDE database is still on the PRIDE XML for conversion of the A.R. J. S. J. J. S. R. Binz P.A. Deutsch E.W. Hermjakob H. Reisinger F. J. J.A. Pizarro A. Creasy D. The data standard for mass proteomics Proteomics. 2012; PubMed Scopus Google Scholar) and L. D. F. J. A. S. Pizarro L. N. Reisinger F. Hermjakob H. Binz P.A. Deutsch E.W. standard for mass spectrometry Proteomics. PubMed Scopus Google Scholar) formats, the standard formats for mass spectrometry data, and are provided by the PRIDE Converter This is an we are currently and support in PRIDE. However, we will to support PRIDE XML as a submission at for the for practical it will some reliable and to for the new data become available for the search and are several existing that PRIDE XML files that we to to support as at to are by the This is the for the Proteomics and J. F. The software an for and of proteomics Res. 2009; PubMed Scopus Google Scholar), of the limitations of the PRIDE XML format is the support for protein can be not in an way the for all the are in the PRIDE XML However, the user can still to the by using the PRIDE XML use that it is not supported by the PRIDE XML format is the for the of several can this information using a of several more the in the and of proteomics a major we the PRIDE Converter 2 to be a major in the capture and of proteomics data, and a key in data submissions to the ProteomeXchange We would to for input during the starting of the files
Côté et al. (2012) studied this question.