PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 13, 2026Cochrane Evidence Synthesis and Methods2 citationsOpen Access

Batch Size Effects on Mid‐2025 State‐of‐the‐Art Large Language Model Performance in Automated Title and Abstract Screening

View Full Paper
PFPetter FagerbergOSOscar SallanderKPKim Vikhe Patil

Key Points

  • To assess how batch size affects the performance of various state-of-the-art large language models in screening references for systematic reviews.
  • Utilized a dataset of 790 references from a Cochrane Review on stem cell treatment.
  • Submitted batches of references ranging from 1 to 790 to four LLMs via public APIs.
  • Measured performance using sensitivity and specificity across different batch sizes with 10 repeated runs for validation.
  • Gemini 2.5 Pro processed all 790 references effectively, while GPT-5 failed at batch sizes ≥400.
  • Sensitivity of GPT-5 mini dropped significantly at higher batch sizes, from 0.88 at batch 200 to 0.48 at batch 400.
  • At a practical batch size of 100, Gemini 2.5 Pro achieved perfect sensitivity (1.00), while GPT-5 had the highest specificity (0.98).

Abstract

ABSTRACT Background Manual abstract screening is a primary bottleneck in evidence synthesis. Emerging evidence suggests that large language models (LLMs) can automate this task, but their performance when processing multiple references simultaneously in “batches” is uncertain. Objectives To evaluate the classification performance of four state‐of‐the‐art LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT‐5, and GPT‐5 mini) in predicting reference eligibility across a wide range of batch sizes for a systematic review of randomized controlled trials. Methods We used a gold‐standard dataset of 790 references (93 considered relevant) from a published Cochrane Review on stem cell treatment for acute myocardial infarction. Using the public APIs for each model, batches of 1 to 790 references were submitted to classify each as “Include” or “Exclude.” Performance was assessed using sensitivity and specificity, with internal validation conducted through 10 repeated runs for each model‐batch combination. Results Gemini 2.5 Pro was the most robust model, successfully processing the full 790‐reference batch. In contrast, GPT‐5 failed at batches ≥400, while GPT‐5 mini and Gemini 2.5 Flash failed at the 790‐reference batch. Overall, all models demonstrated strong performance within their operational ranges, with two notable exceptions: Gemini 2.5 Flash showed low initial sensitivity at batch 1, and GPT‐5 mini's sensitivity degraded at higher batch sizes (from 0.88 at batch 200 to 0.48 at batch 400). At a practical batch size of 100, Gemini 2.5 Pro achieved the highest sensitivity (1.00, 95% CI 1.00–1.00), whereas GPT‐5 delivered the highest specificity (0.98, 95% CI 0.98–0.98). Conclusion State‐of‐the‐art LLMs can effectively screen multiple abstracts per prompt, moving beyond inefficient single‐reference processing. However, performance is model‐dependent, revealing trade‐offs between sensitivity and specificity. Therefore, batch size optimization and strategic model selection are important parameters for successful implementation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fagerberg et al. (2026) studied this question.

synapsesocial.com/papers/69dc887f3afacbeac03ea589https://doi.org/10.1002/cesm.70082
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Scaling the Prompt: How Batch Size Shapes Performance of Mid-2025 State-of-the-Art LLMs in Automated Title-and-Abstract Screening2025
  2. 2Compact large language models for title and abstract screening in systematic reviews: An assessment of feasibility, accuracy, and workload reduction2025
  3. 3Large language models for abstract screening in systematic- and scoping reviews: A diagnostic test accuracy study2024 · 3 citations
  4. 4Comparative evaluation of large language models for guideline-compliant abstract generation and readability in dental research: an experimental comparative study2026
  5. 5Evaluating the Effectiveness of Large Language Models in Abstract Screening: A Comparative Analysis2024 · 5 citations