Mass spectrometry-based proteomics usually relies on canonical protein sequences, yet proteins often exist in multiple forms due to alternative ORFs, splicing, or genetic variants. These sequence differences can influence protein function and disease outcomes, but they remain largely invisible to conventional workflows. AliceDB addresses this by integrating UniProt and ClinVar into ∼7 million human protein entries and generating customized, project-specific FASTA files that include variants. This is especially valuable for de novo sequencing, where many peptides cannot be assigned to canonical proteins; with AliceDB, more sequences can be mapped, increasing biological interpretability. We present the complete AliceDB workflow, from database construction to variant-aware peptide identification and downstream statistical analysis. As a proof of concept, we analyzed 24 follicular fluid samples from 12 patients to search for molecular markers of oocyte competence. Incorporating sequence variants expanded peptide identifications by ∼5% over canonical searches and revealed alterations associated with fertilization outcomes. Although demonstrated in reproductive biology, AliceDB is broadly applicable to proteomics and biomarker discovery, where even small gains in identification can uncover clinically relevant signals. By systematically incorporating natural protein variants, AliceDB exposes hidden layers of proteomic diversity, supporting both precision medicine and functional biology.
Mruk et al. (2026) studied this question.