Reproducibility serves as the foundation upon which robust scientific discoveries are built. When designed following best-practices, analyzes in computational biology can achieve this ideal, facilitating straightforward interpretation and generation of hypotheses. The recent surge in big data generated from biomedical measurement techniques, as well as the need to register their dimensions of information, necessitates the development of novel algorithms. Additionally, to facilitate quick adoption and ease of use while maintaining full reproducibility and high analyzes standards, workflows and modality-specific resource databases must be created and disseminated. In this thesis, I employed the workflow management system Snakemake, in conjunction with Datavzrd for data compilation and distribution, to address the need for data analysis pipelines and reference databases. The development and application of these workflows and resources resulted in three first author publications and culminated in this cumulative thesis. Spectral libraries derived from LC-MS/MS proteome data are invaluable resources that support the design of targeted mass spectrometry experiments and help identify low-abundant proteins. In addition, they can serve as a database for DIA measurement strategies to deconvolute the highly complex spectra. Despite recent efforts to generate cell-specific spectral libraries, there is still a lack of comprehensive information for many immune cells. To improve this, we developed the Spectral Library of Immune Cells (SpLICe), which contains detailed proteomic measurements of macrophages, dendritic cells, B-cells and CD4 and CD8 T-cells. This reference database comprises almost 9,000 protein groups and more than 110,000 proteotypic peptides across all cell types. In addition to the typical contents of a spectral library, we also included post-translational modification site maps and enrichment terms from the Gene Ontology and Reactome databases for each protein. Furthermore, we compiled a list of all detected peptides for every protein and assigned each one a peptide score that quantifies its suitability for targeted mass spectrometry analysis. Finally, we utilized Datavzrd to assemble all the information into an interactive, visual, and server-free HTML file that is ready to be used for designing mass spectrometry experiments that target the featured immune cell types. Moreover, I developed a quantitative proteomics data analysis workflow. We used this pipeline to analyze the proteomic changes in the forebrains of 3-week-old mice that have a MCT8/OATP1C1 double-knockout. A lack of MCT8 transporter function in humans results in the Allan-Herndon-Dudley syndrome, a developmental disorder associated with impaired locomotor function and cognition, as well as delayed myelination. Our analysis revealed that double-knockout mice display reduced levels of proteins related to myelination, oligodendrocytes, and neurogenesis, which aligns with findings in previous studies. In addition, we compared our results with open-access RNA sequencing data and found that both modalities display significantly reduced expression of Pde10a and Pvalb, both of which have been linked to locomotor disorders. Through the analysis of proteomic data, we identified several altered proteins that may help elucidate the mechanisms behind the complications associated with dysfunctional MCT8 in humans. We also applied this pipeline to a mouse model, in which we compared the proteome of wild type mice with mice that were injected with Stx and LPS and mice that were additionally treated with Etanercept. Etanercept competitively inhibits TNF-alpha which we hypothesized to play a key role in the brain pathologies that can develop during HUS caused by Shiga toxin-producing Escherichia coli. Previous studies have documented that the mortality rate of the HUS is highly correlated with the development of brain pathologies. Our proteomic data has demonstrated that, for molecules belonging to the complement cascade, Etanercept-treated mice display an expression similar to wild type mice, reversing the upregulation observed by Stx/LPS injection. This is consistent with our findings in macroscopy of ameliorated levels of angiopathy. To investigate the microglia morphology in this study, I developed a machine learning-based, open-access, and reusable Snakemake segmentation and immune cell shape analysis workflow. Using this pipeline, we have shown that Etanercept mitigates the morphological changes in microglia induced by injecting Stx and LPS, thereby linking TNF-alpha activity with microglia activation. In summary, the studies in this thesis demonstrate the essential role of analysis workflows and resources for the reproducible analysis of biomedical data from all fields, robustly identifying potential biomarkers while being flexible enough to be reused and extended by the scientific community.
Devon Siemes (Wed,) studied this question.