PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 21, 2026Cell Reports Medicine2 citationsOpen Access

Benchmarking large language models for predictive modeling in biomedical research with a focus on reproductive health

View Full Paper
RSReuben D. SarwalVTVictor TarcaCDClaire Dubin

Key Points

  • The study aims to evaluate the performance of large language models in predictive tasks relevant to reproductive health.
  • Assess performance of eight large language models across four predictive tasks from DREAM challenges.
  • Use task descriptions, data locations, and target outcomes to prompt LLMs for code generation.
  • Run LLM-generated code to fit prediction models and determine accuracy on test sets.
  • Four LLMs successfully complete at least one task.
  • R code generation is more successful than Python code generation (14/16 vs. 7/16).
  • OpenAI's o3-mini-high performs best, completing 7/8 tasks.
  • Top LLM-generated models match or exceed median-participating team performance, with one task surpassing the top team (p = 0.02).

Abstract

Large language models (LLMs) are increasingly used for code generation and data analysis. This study assesses LLM performance across four predictive tasks from three DREAM challenges: gestational age regression from transcriptomics and DNA methylation and classification of preterm birth and early preterm birth from microbiome data. We prompt LLMs with task descriptions, data locations, and target outcomes and then run LLM-generated code to fit prediction models and determine accuracy on test sets. Among the eight LLMs tested, o3-mini-high, 4o, DeepseekR1, and Gemini 2.0 can complete at least one task. R code generation is more successful (14/16) than Python (7/16). OpenAI's o3-mini-high outperforms others, completing 7/8 tasks. Test set performance of the top LLM-generated models matches or exceeds the median-participating team for all four tasks and surpasses the top-performing team for one task (p = 0.02). These findings underscore the potential of LLMs to democratize predictive modeling in omics and increase research output.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sarwal et al. (2026) studied this question.

synapsesocial.com/papers/69994a7f873532290d01ef89https://doi.org/10.1016/j.xcrm.2026.102594
Ask AI
Helpful
Bookmark
Share
View Full Paper