PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 2025The Lancet Digital Health71 citationsOpen Access

Importance of sample size on the quality and utility of AI-based prediction models for healthcare

RRRichard D RileyJEJoie EnsorKSKym I E Snell

Key Points

Key points are not available for this paper at this time.

Abstract

Rigorous study design and analytical standards are required to generate reliable findings in healthcare from artificial intelligence (AI) research. One crucial but often overlooked aspect is the determination of appropriate sample sizes for studies developing AI-based prediction models for individual diagnosis or prognosis. Specifically, the number of participants and outcome events required in datasets for model training and evaluation remains inadequately addressed. Most AI studies do not provide a rationale for their chosen sample sizes and frequently rely on datasets that are inadequate for training or evaluating a clinical prediction model. Among the ten principles of Good Machine Learning Practice established by the US Food and Drug Administration, the UK Medicines and Healthcare products Regulatory Agency, and Health Canada, guidance on sample size is directly relevant to at least three principles. To reinforce this recommendation, we outline seven reasons why inadequate sample size negatively affects model training, evaluation, and performance. Using a range of examples, we illustrate these issues and discuss the potentially harmful consequences for patient care and clinical adoption. Additionally, we address challenges associated with increasing sample sizes in AI research and highlight existing approaches and software for calculating the minimum sample sizes required for model training and evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Riley et al. (2025) studied this question.

synapsesocial.com/papers/6a0662fd3f8bf83a443dd9c0https://doi.org/10.1016/j.landig.2025.01.013
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Ethical Machine Learning in Healthcare2021 · 541 citations
  2. 2Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations2021 · 893 citations
  3. 3A decomposition of Fisher's information to inform sample size for developing fair and precise clinical prediction models -- part 1: binary outcomes2024 · 2 citations
  4. 4Minimum sample size for external validation of a clinical prediction model with a continuous outcome2020 · 201 citations
  5. 5External validation of clinical prediction models: simulation-based sample size calculations were more reliable than rules-of-thumb2021 · 123 citations