Machine learning is increasingly used in entrepreneurship analytics to predict startup outcomes, frequently reporting accuracy above 0.90, yet whether such performance reflects a genuine ex-ante signal or methodological artifact remains unclear and consequential for investors, accelerators, and innovation-policy agencies. This study evaluates startup-outcome classification under a leakage-controlled, calibration-first protocol using two Crunchbase-derived datasets (66,368 firms; a 923-firm engineered-feature set), three success constructs, and three model families under five-fold stratified cross-validation. Removing outcome-correlated, survivorship-accumulating features lowers the area under the receiver operating characteristic curve by 0.05 to 0.09 on the large dataset, with every paired 95% confidence interval excluding zero, and by 0.19 on the engineered dataset; an independent study on the same 923-firm data without leakage control reports 88.1% accuracy. The leakage-controlled performance level is modest (0.66 to 0.77). Calibration rankings diverge from discrimination rankings: gradient boosting is well calibrated (expected calibration error of 0.006 to 0.024), whereas logistic regression shows large calibration error on imbalanced constructs (largely an artifact of class weighting rather than an intrinsic model property); post hoc isotonic recalibration then removes most of the error. The contribution is a reusable evaluation protocol for entrepreneurship analytics. Findings are associational and specific to the analyzed samples.
Khamphukun et al. (Mon,) studied this question.