PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 2026Journal of the American Medical Informatics Association3 citations

Is one run enough? Reproducibility of flagship large language models across temperature and reasoning settings in biomedical text processing

View Full Paper
PWPaul WindischCKCarole KoechliFDFabio Dennstädt

Key Points

  • This research aims to evaluate the run-to-run reproducibility of large language models in classifying biomedical trials.
  • Utilized 250 trial abstracts labeled for endpoint success.
  • Evaluated models across various thinking levels and temperature settings.
  • Each experimental setting was executed three times.
  • High reproducibility observed for both models with κ values near 1.0.
  • F1 scores remained stable between 0.955 and 0.971.
  • Minor gains were noted from majority voting among runs.

Abstract

Abstract Background To quantify run-to-run reproducibility of Gemini 3 Flash Preview and GPT-5.2 for trial-success classification across temperature and reasoning/thinking settings and determine whether single-run reporting suffices. Materials and Methods We utilized 250 trial abstracts labeled based on primary endpoint success. We evaluated Gemini across thinking levels (minimal, low, medium, high) and temperatures 0.0-2.0 and GPT-5.2 across reasoning-effort levels (none to x-high) with an additional temperature sweep when reasoning was disabled. Each setting was run 3 times. Results Reproducibility was high for Gemini (κ = 0.942-1.000; invalid outputs 0%-1.5%) and GPT-5.2 (κ = 0.984-0.995; no invalid outputs). F1 remained stable (mean/majority vote 0.955-0.971), with marginal gains from majority voting. Conclusion For binary biomedical classification with tightly constrained outputs, both models were reproducible across decoding and reasoning settings, suggesting single runs are often sufficient, with minimal replication as a practical stability check.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Windisch et al. (2026) studied this question.

synapsesocial.com/papers/69ccb5f716edfba7beb87b8ehttps://doi.org/10.1093/jamia/ocag039
Ask AI
Helpful
Bookmark
Share
View Full Paper