PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 22, 2026Information0 citationsOpen Access

A Leakage-Controlled, Calibration-First Evaluation of Machine Learning Models for Startup-Outcome Prediction: Evidence from Crunchbase

View Full Paper
RKRatchaneekorn KhamphukunWNWarawut Narkbunnum

Key Points

  • This study aims to assess the effectiveness of machine learning models in predicting startup outcomes while addressing methodological artifacts.
  • Evaluated startup outcomes using a leakage-controlled, calibration-first protocol.
  • Utilized two Crunchbase datasets with 66,368 firms and a 923-firm engineered-feature set.
  • Applied five-fold stratified cross-validation across three model families.
  • Leakage-controlled models achieve an area under the curve of 0.66 to 0.77, indicating modest performance.
  • Removal of outcome-correlated features results in accuracy reductions of 0.05 to 0.09 on large dataset and 0.19 on engineered dataset.
  • Gradient boosting shows low calibration error, while logistic regression requires recalibration to correct large errors.

Abstract

Machine learning is increasingly used in entrepreneurship analytics to predict startup outcomes, frequently reporting accuracy above 0.90, yet whether such performance reflects a genuine ex-ante signal or methodological artifact remains unclear and consequential for investors, accelerators, and innovation-policy agencies. This study evaluates startup-outcome classification under a leakage-controlled, calibration-first protocol using two Crunchbase-derived datasets (66,368 firms; a 923-firm engineered-feature set), three success constructs, and three model families under five-fold stratified cross-validation. Removing outcome-correlated, survivorship-accumulating features lowers the area under the receiver operating characteristic curve by 0.05 to 0.09 on the large dataset, with every paired 95% confidence interval excluding zero, and by 0.19 on the engineered dataset; an independent study on the same 923-firm data without leakage control reports 88.1% accuracy. The leakage-controlled performance level is modest (0.66 to 0.77). Calibration rankings diverge from discrimination rankings: gradient boosting is well calibrated (expected calibration error of 0.006 to 0.024), whereas logistic regression shows large calibration error on imbalanced constructs (largely an artifact of class weighting rather than an intrinsic model property); post hoc isotonic recalibration then removes most of the error. The contribution is a reusable evaluation protocol for entrepreneurship analytics. Findings are associational and specific to the analyzed samples.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khamphukun et al. (2026) studied this question.

synapsesocial.com/papers/6a605dd44163e025518d7badhttps://doi.org/10.3390/info17070702
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Predicting startup success using two bias-free machine learning: resolving data imbalance using generative adversarial networks2024 · 15 citations
  2. 2Start-up Acquisition Status Prediction Using Machine Learning2023 · 1 citations
  3. 3To Explain or to Predict?2010 · 2,475 citations
  4. 4Predicting Startup Success Using Tree-Based Machine Learning Algorithms2024 · 5 citations
  5. 5Assessing Algorithmic Fairness with Unobserved Protected Class Using Data Combination2021 · 85 citations