The open XML format mzML, used for representation of MS data, is pivotal for the development of platform-independent MS analysis software. Although conversion from vendor formats to mzML must take place on a platform on which the vendor libraries are available (i.e. Windows), once mzML files have been generated, they can be used on any platform. However, the mzML format has turned out to be less efficient than vendor formats. In many cases, the naïve mzML representation is fourfold or even up to 18-fold larger compared with the original vendor file. In disk I/O limited setups, a larger data file also leads to longer processing times, which is a problem given the data production rates of modern mass spectrometers. In an attempt to reduce this problem, we here present a family of numerical compression algorithms called MS-Numpress, intended for efficient compression of MS data. To facilitate ease of adoption, the algorithms target the binary data in the mzML standard, and support in main proteomics tools is already available. Using a test set of 10 representative MS data files we demonstrate typical file size decreases of 90% when combined with traditional compression, as well as read time decreases of up to 50%. It is envisaged that these improvements will be beneficial for data handling within the MS community. The open XML format mzML, used for representation of MS data, is pivotal for the development of platform-independent MS analysis software. Although conversion from vendor formats to mzML must take place on a platform on which the vendor libraries are available (i.e. Windows), once mzML files have been generated, they can be used on any platform. However, the mzML format has turned out to be less efficient than vendor formats. In many cases, the naïve mzML representation is fourfold or even up to 18-fold larger compared with the original vendor file. In disk I/O limited setups, a larger data file also leads to longer processing times, which is a problem given the data production rates of modern mass spectrometers. In an attempt to reduce this problem, we here present a family of numerical compression algorithms called MS-Numpress, intended for efficient compression of MS data. To facilitate ease of adoption, the algorithms target the binary data in the mzML standard, and support in main proteomics tools is already available. Using a test set of 10 representative MS data files we demonstrate typical file size decreases of 90% when combined with traditional compression, as well as read time decreases of up to 50%. It is envisaged that these improvements will be beneficial for data handling within the MS community. Open XML formats for representation of MS data have been developed by the proteomics community to facilitate exchange and vendor neutral analysis of mass spectrometry data. Initially two formats, mzXML (1.Pedrioli P.G.A. Eng J.K. Hubley R. Vogelzang M. Deutsch E.W. Raught B. Pratt B. Nilsson E. Angeletti R.H. Apweiler R. Cheung K. Costello C.E. Hermjakob H. Huang S. Julian R.K. Kapp E. McComb M.E. Oliver S.G. Omenn G. Paton N.W. Simpson R. Smith R. Taylor C.F. Zhu W. Aebersold R. A common open representation of mass spectrometry data and its application to proteomics research.Nat. Biotechnol. 2004; 22: 1459-1466Crossref PubMed Scopus (652) Google Scholar) and mzData (http://psidev.info/), existed in parallel, until these formats were merged into the single standard format mzML (2.Martens L. Chambers M. Sturm M. Kessner D. Levander F. Shofstahl J. Tang W.H. Römpp A. Neumann S. Pizarro A.D. Montecchi-Palazzi L. Tasman N. Coleman M. Reisinger F. Souda P. Hermjakob H. Binz P.-A. Deutsch E.W. mzML–a community standard for mass spectrometry data.Mol. Cell. Proteomics. 2011; 10Abstract Full Text Full Text PDF PubMed Scopus (452) Google Scholar). The mzML format has been adopted widely by the proteomics community and is supported by many data processing tools. However, although successfully used in many pipelines, the mzML format has not reached its full usage potential, mainly because of large file sizes in comparison to the raw vendor formats. The file size problem has become more marked with the introduction of recent high-resolution high-frequency mass spectrometers. As an example, a raw data file from a data-independent acquisition experiment using an AB SCIEX TripleTOF resulted in a vendor format data file of 2.5 GB. Conversion of this file to standard mzML resulted in a 46.7 GB file, with a conversion time of about 12 min on a desktop computer (later called high-end I, Fig. 1A) dedicated to the conversion process. If the file is compressed using gzip to lower the storage footprint, the size drops to 21.6 GB, but the conversion now takes 2 h instead. This (extreme) example pinpoints the need for increased efficiency in the standardized representation of MS data. Across the community, there is little previous work done on compression of MS data. In two technical papers (3.Miguel A.C. Keane J.F. Whiteaker J. Zhang H. Paulovich A.G. Compression of LC/MS Proteomic Data.Proc. 19th IEEE Symp. Comput.-Based Med. Syst. CBMS06. 2006; Crossref Scopus (2) Google Scholar, 4.Miguel A.C. Kearney-Fischer M. Keane J.F. Whiteaker J. Feng L.-C. Paulovich A.G. Near-lossless compression of mass spectra for proteomics.Acoust. Speech Signal Process. 2007 ICASSP 2007 IEEE Int. Conf. 2007; 1: I-369-I-372Google Scholar), Miguel et al. describe lossless and near-lossless compression methods for QTOF data, achieving compression factors above 10. No measurements are provided on the compression time however, and the algorithms are benchmarked on a very small set of data files. Blanckenburg et al. describe a lossy compression technique for Fourier transform ion cyclotron resonance data (5.Blanckenburg B. Burgt Y.E.M. Deelder A.M. Palmblad M. “Lossless” compression of high resolution mass spectra of small molecules.Metabolomics. 2010; 6: 335-340Crossref PubMed Scopus (1) Google Scholar), where known nonmetabolite data points are discarded. Outside the MS community, potential benefits could come from recent work in the numerical computation field (6.Engelson V. Fritzson D. Fritzson P. Lossless Compression of High-volume Numerical Data from Simulations.Data Compression Conf. 2000; : 574-586Google Scholar, 7.Ratanaworabhan P. Ke J. Burtscher M. Fast lossless compression of scientific floating-point data.Data Compression Conf. 2006 DCC 2006 Proc. 2006; 1: 133-142Crossref Scopus (109) Google Scholar), were many data types are similar to MS data in terms of precision and smoothness. Nevertheless, perhaps the most relevant recent advance is the emergence of the mz5 format (8.Wilhelm M. Kirchner M. Steen J.A.J. Steen H. mz5: space- and time-efficient storage of mass spectrometry data sets.Mol. Cell. Proteomics. 2012; 11Abstract Full Text Full Text PDF PubMed Scopus (45) Google Scholar), which yields performance increases via a binary representation and optimized libraries, as well as some regular data compression. Although this format, based on the open HDF5 standard (The HDF Group, Champaign, IL, USA), is an efficient representation of mzML files, it suffers from the fact that the files are not readable without native libraries or specialized software, which, to some extent, has hampered its uptake. Also, while mz5 can be “lossless,” default compression implies removal of zero intensity scans, which means the original data cannot be reconstructed, and some algorithms require zero intensity scans for correct functioning. The standard XML representation used in mzML can be easily viewed as text on any operating system, and it is relatively easy to write a parser in any programming language. We thus sought to overcome the mzML efficiency shortcomings by introducing better compression of the binary data found in mzML files while still leaving the metadata in XML format, and propose such an extension to the format here. Furthermore, we envisage that decompression of this binary data should be easy to incorporate into software tools via permissively licensed stand-alone source code files for C++ and Java, which do not require any external dependences. We here also exemplify the facility of usage by implementing support in several popular tools for proteomics data analysis. To efficiently compress the three main types of binary data present in mzML files: (1) mass to charge ratios, (2) ion counts, and (3) retentions times, we have developed three new near-lossless compression algorithms, while ensuring for each data type that precision losses are well below the precision of the most advanced mass spectrometers of today. The Numpress Linear Prediction Compression algorithm (hereafter called numLin, relative error < 2e-9) takes advantage of the linearly increasing values in m/z and retention time data, and is optimized for high-resolution m/z data. Ion count data does not linearly increase but requires less stored precision because of the lower instrument precision, and Numpress Short Logged Float (numSlof 1The abbreviations used are: numSlof, numpress short logged float; XML, extensible markup language; DDA, data-dependent acquisition; SRM, selected reaction monitoring; DIA, data-independent acquisition; numLin, numpress linear prediction; numPic, numpress positive integer count; numSafe, numpress linear prediction transformation (lossless); numAll, combination of numLin and numSlof compression; mz5zlib, mz5 with zlib compression; SSD, solid-state drive; MGF, mascot generic format. , relative error < 2e-4) is optimized for this data type. We also developed a second ion count compression (Numpress Positive Integer Count, numPic) and a lossless transformation (Numpress Linear Prediction Transformation, numSafe), which are not used further here, but are presented in the supplementary materials. Although the least significant of the 16 double-precision decimals are lost in the first conversion to the compressed format for all the algorithms, compression and decompression after this does not incur further losses. To maximize speed, the algorithms are highly local in memory and only need a single traversal of the data. For a complete description of the algorithms we refer to Supplemental Methods, and to the reference implementations in Java and C++, found at https://github.com/ms-numpress/ms-numpress under the Apache 2.0 license. To compare MS-Numpress to current alternatives for storing mzML data, we extensively evaluated size, write time, and read time of available compression schemes on a varied set of data files using different computers (Fig. 1). For this we constructed a test set of 10 MS data files from different vendors, instruments, and experiment types (Fig. 1D and supplemental Table S1). The test set files included data-dependent acquisition (DDA), selected reaction monitoring (SRM) and data-independent (DIA/SWATH) acquisition modes, and both simple and complex samples, giving a heterogeneous set of distributions of MS1 and MS2 spectrum data and chromatogram binary data arrays of different lengths (supplemental Fig. S1). These files were converted to mzML, imzML (9.Römpp A. Schramm T. Hester A. Klinkert I. Both J.-P. Heeren R.M.A. Stöckli M. Spengler B. imzML: Imaging Mass Spectrometry Markup Language: A common data format for mass spectrometry imaging.Methods Mol. Biol. 2011; 696: 205-224Crossref PubMed Scopus (54) Google Scholar) and mz5 (8.Wilhelm M. Kirchner M. Steen J.A.J. Steen H. mz5: space- and time-efficient storage of mass spectrometry data sets.Mol. Cell. Proteomics. 2012; 11Abstract Full Text Full Text PDF PubMed Scopus (45) Google Scholar), both without compression, using zlib compression, and using gzip of the entire file. Files were also compressed in multiple different setups using MS-Numpress compressions, resulting in a total of 18 tested compression schemes (Fig. 1C and supplemental Table S2). To avoid clutter, minor results are left out here, and readers interested in imzML-data or individual Numpress results are referred to the supplementary material. The different compression schemes were compared based on file size, read time and write time (Fig. 1B). Benchmarking was performed on four dedicated desktop computers of varying capacity (Fig. 1A), using a custom msconvert (10.Chambers M.C. Maclean B. Burke R. Amodei D. Ruderman D.L. Neumann S. Gatto L. Fischer B. Pratt B. Egertson J. Hoff K. Kessner D. Tasman N. Shulman N. Frewen B. Baker T.A. Brusniak M.-Y. Paulse C. Creasy D. Flashner L. Kani K. Moulding C. Seymour S.L. Nuwaysir L.M. Lefebvre B. Kuhlmann F. Roark J. Rainer P. Detlev S. Hemenway T. Huhmer A. Langridge J. Connolly B. Chadick T. Holly K. Eckels J. Deutsch E.W. Moritz R.L. Katz J.E. Agus D.B. MacCoss M. Tabb D.L. Mallick P. A cross-platform toolkit for mass spectrometry and proteomics.Nat. Biotechnol. 2012; 30: 918-920Crossref PubMed Scopus (1775) Google Scholar) build, and timed using a script written in Python. Write time was measured as the total time for an msconvert conversion from the vendor raw format. Because this includes the vendor read time it gives a constant offset, but this constant is in general small compared with the write time. For read benchmarking a custom program was made, that reads files using the ProteoWizard (10.Chambers M.C. Maclean B. Burke R. Amodei D. Ruderman D.L. Neumann S. Gatto L. Fischer B. Pratt B. Egertson J. Hoff K. Kessner D. Tasman N. Shulman N. Frewen B. Baker T.A. Brusniak M.-Y. Paulse C. Creasy D. Flashner L. Kani K. Moulding C. Seymour S.L. Nuwaysir L.M. Lefebvre B. Kuhlmann F. Roark J. Rainer P. Detlev S. Hemenway T. Huhmer A. Langridge J. Connolly B. Chadick T. Holly K. Eckels J. Deutsch E.W. Moritz R.L. Katz J.E. Agus D.B. MacCoss M. Tabb D.L. Mallick P. A cross-platform toolkit for mass spectrometry and proteomics.Nat. Biotechnol. 2012; 30: 918-920Crossref PubMed Scopus (1775) Google Scholar) API. To ensure that all data is read, this program explicitly reads all binary values in the spectra and chromatograms found in the file. Test files, results, and program binaries are available at the Swestore repository (http://webdav.swegrid.se/snic/bils/lu_proteomics/pub). The use of near-lossless compression introduces the question of whether one can be sure that no analytically relevant data is lost. We measured the relative errors for the compressed versions of all the files in the test set (supplemental Tables S3 to S6), and found relative errors to be smaller than 2e-9 (0.002 ppm) for numLin compressed m/z data and smaller than 2e-4 (0.02%) for numSlof compressed ion counts. To validate that the small errors introduced by Numpress compression do not have adverse effect on common proteomics analyses, we converted two Orbitrap DDA LC-MS/MS mzML files to compressed (combination of numLin and numSlof) versions and back to uncompressed mzML again, and compared analysis results obtained from the doubly converted files to those obtained using the original files. MS/MS identification using Mascot after extraction of MGF (Mascot Generic Format) peak lists yielded identical lists of identified peptides at a 1% peptide to spectrum match (PSM) false discovery rate (FDR, Supplementary data). Extraction of features from the MS1 data using msInspect M. M. M. M. T. P. D. Eng J. R. C. J. D. Whiteaker J. Paulovich A. M. A of algorithms for the analysis of complex using high-resolution 2006; 22: PubMed Scopus Google Scholar) yielded the lists of with only in the least significant decimals of some m/z values and lists of for the peptides that were identified using the lists the with a relative intensity introduced by the that the Numpress compression schemes could be used for proteomics We found that to file size, a combination of numLin for m/z or retention time data and numSlof for ion count data with of the entire file, was the most in terms of file This yielded an file size of compared with standard mzML all 10 test set files (Fig. with longer write (Fig. but read (Fig. on all tested computers and files (supplemental Table This format is also the size of the binary mz5 with zlib compression and also smaller than all vendor formats for AB files (supplemental Fig. and Table In read the text based mzML formats cannot with the binary although the is small in the files (Fig. The of disk I/O compression are the most in the files on the lower performance computers (supplemental S3 and where the alternatives up to read of (supplemental Fig. and write four of (supplemental Fig. of the test computers were with which are to open up the disk I/O and the these the of with increasing with file size (Fig. and the from the disk I/O are for the schemes mzML, mz5zlib, and (supplemental and compression of individual data arrays also minor of the write the does not write at all (Fig. read by on compared with standard mzML (supplemental Table further read by from and this have a of for this are (1) the file of mz5 read (2) large of while the mzML is or (3) is (3) is not while a format, both (1) and (2) could be in optimized As the MS-Numpress compression high of compression, which was we set out to the technique as of several proteomics in to ensure easy by the proteomics community. The in ProteoWizard (10.Chambers M.C. Maclean B. Burke R. Amodei D. Ruderman D.L. Neumann S. Gatto L. Fischer B. Pratt B. Egertson J. Hoff K. Kessner D. Tasman N. Shulman N. Frewen B. Baker T.A. Brusniak M.-Y. Paulse C. Creasy D. Flashner L. Kani K. Moulding C. Seymour S.L. Nuwaysir L.M. Lefebvre B. Kuhlmann F. Roark J. Rainer P. Detlev S. Hemenway T. Huhmer A. Langridge J. Connolly B. Chadick T. Holly K. Eckels J. Deutsch E.W. Moritz R.L. Katz J.E. Agus D.B. MacCoss M. Tabb D.L. Mallick P. A cross-platform toolkit for mass spectrometry and proteomics.Nat. Biotechnol. 2012; 30: 918-920Crossref PubMed Scopus (1775) Google Scholar) conversion to the format from all mass raw data formats, and read to tools that use the ProteoWizard for files, for example D.L. Chambers M.C. highly mass peptide identification by 2007; 6: PubMed Scopus Google Scholar) and B. Shulman N. Chambers M. Frewen B. R. Tabb D.L. MacCoss an open source for and proteomics 2010; PubMed Scopus Google Scholar). Compression should be for that use high-resolution data for because of the large data files this and we support for of mzML files with numpress compressed binaries in K. C. E. N. Sturm M. proteomics 2007; PubMed Scopus Google Scholar), msInspect M. M. M. M. T. P. D. Eng J. R. C. J. D. Whiteaker J. Paulovich A. M. A of algorithms for the analysis of complex using high-resolution 2006; 22: PubMed Scopus Google Scholar), and the J. G. K. Levander F. The software an extensible platform for and analysis of proteomics PubMed Scopus Google Scholar, M. A. K. E. S. Levander F. algorithm for Cell. Full Text Full Text PDF PubMed Scopus Google Scholar). The in also implies that the complete MS-Numpress compression and decompression algorithms are available in the because of the recent of the complete Aebersold R. L. A to the algorithm PubMed Scopus Google Scholar). The Reisinger F. L. an Java for mzML, the standard for MS 2010; PubMed Scopus Google Scholar) has also been with MS-Numpress read and write support for numLin, numSlof, and The is in mass spectrometry software J. R. B. N. L. Reisinger F. A. D. H. Hermjakob H. The 2 an of tools to facilitate data to the and the Cell. Proteomics. 2012; Full Text Full Text PDF PubMed Scopus Google Scholar) and an open source for the analysis of proteomics data will now support MS-Numpress compressed data when the for mzML with MS-Numpress compression was in the tools of the A. Eng J. Zhang N. Aebersold R. A proteomics MS/MS analysis platform open XML file Syst. Biol. 1: PubMed Scopus Google Scholar, E.W. L. D. T. H. Tasman N. Nilsson E. Pratt B. B. Eng J.K. D.B. Aebersold R. A of the 2010; PubMed Scopus Google Scholar). was also in the R. with mass 2004; PubMed Scopus Google Scholar). we support for MS-Numpress in the J. C. S. K. P. J. Levander F. selected reaction monitoring software for 2012; PubMed Scopus Google Scholar) for The main of using the new mzML format is where small file sizes are of for example in data the file size is also of in or for example by et al. for a proteomics for 2004; PubMed Scopus Google Scholar). Because of the file size simple peak lists formats are still used extensively for MS/MS and the more complete file provided by mzML have mainly been used for data and However, with the increased of MS/MS spectra with modern mass spectrometers it is envisaged that compressed data formats could become more widely used also for MS/MS As an example, we a raw data file from a LC-MS/MS analysis of a performed on an Orbitrap for a recent from the A. The Cell. Proteomics. Full Text Full Text PDF PubMed Scopus Google Scholar). Conversion of the file to Mascot Generic using ProteoWizard and no yielded a file size of with only the MS/MS data ProteoWizard conversion of the original file using identical to mzML with compression resulted in a file, while still the The file size to only the MS/MS data is in the mzML file. an MS/MS in resulted in more peptide at a 1% using the mzML file than using the MGF with a mass of because of a precision in the mass in the mzML file, that there is no in using the compressed format for MS/MS results demonstrate the of some very simple to the mzML format with to disk and handling time. are algorithms that could further on the of compression or handling times, but for standard formats we it is to simple and to both the of in and the of in the We further support for the new algorithms in several and thus to the proteomics community. MS-Numpress will also be evaluated the for in the mzML We that this work also data compression algorithm which, in the will to an to the mzML standard to data handling for all mzML We for the C++ MS-Numpress and to for the QTOF data file. with files
No takes yet. Share an insight, caveat, or question.
Teleman et al. (2014) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: