Key points are not available for this paper at this time.
Background Preterm birth (PTB, 37 weeks of gestation) remains a major cause of neonatal morbidity and mortality worldwide, with Hispanic/Latino populations markedly underrepresented in microbiome-based studies, particularly in intensive data analytics scenarios. Methods We applied leakage-aware machine learning as a descriptive analytical framework to characterize clinical and vaginal microbiome patterns associated with preterm birth in 43 pregnant Mexican women (110 longitudinal samples, 14 preterm births) recruited from public hospitals in Mexico City. Vaginal microbiome profiles (genus-level 16S rRNA V3-V4 sequencing) were analyzed using centered log-ratio transformation. We evaluated 12 model configurations representing combinations of two algorithms (Random Forest, Elastic Net), three clinical feature selection strategies (minimal DREAM-style adjustment, literature-based comprehensive features, data-driven empirical selection), and two microbiome representations (ANCOM-BC2 differentially abundant taxa, full filtered profiles). Random Forest and Elastic Net models were implemented within a rigorous subject-level nested cross-validation design to prevent data leakage. Model discrimination metrics were interpreted as indicators of internal cohort structure rather than as estimates of clinical predictive performance. Differential abundance analyses were conducted using ANCOM-BC2 both globally and within cross-validation folds to assess feature robustness. Results The best-performing descriptive model (Random Forest with data-driven feature selection and full microbiome) exhibited AUROC 0.813 ± 0.110 , consistent with structured clinical-microbiome patterning within the cohort. Global differential abundance analysis (ANCOM-BC2, adjusted for maternal age and pre-pregnancy BMI) identified Mycoplasma as the only genus achieving FDR-corrected significance (LFC = + 1.004 , q = 0.049 ), with ten additional genera reaching nominal significance ( p 0.05 ). Within-fold feature importance and stability analyses consistently prioritized anthropometric variables (BMI, pre-pregnancy weight) alongside Peptostreptococcus and Mycoplasma , both detected in 100% of cross-validation iterations, indicating relative signal stability despite limited sample size. Conclusions This study illustrates how descriptive, leakage-aware machine learning can organize, prioritize, and interpret clinical and microbiome signals in small, underrepresented cohorts. At this stage, it does not yet present a clinically deployable predictor for preterm birth, but we are working towards this definite goal in the future, with prenatal screening strategies in mind. The observed internal discrimination reflects, in this sense, cohort-specific structure rather than validated predictive performance and establishes a methodological basis for future externally validated classifiers in Latin American populations.
Ruhle et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: