Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
December 8, 2025Blood

Evaluating large language models in real-world hematologic clinical decision-making: Performance, limitations, and clinical implications

View Full Paper
Ask AI
Bookmark
Share

Authors

ACAngela ConsagraJWJiasheng WangGRGustavo Rivero

Discussion

Loading...

Member takes

Overview

This evaluation demonstrates diagnostic accuracy in hematology cases using AI, suggesting caution in implementing these models.

Key Points

  • The aim is to assess the performance of advanced AI language models in real hematology clinical scenarios.
  • Developed a test set of 30 complex clinical cases of myelodysplastic syndromes (MDS).
  • Evaluated AI models on diagnosis per WHO criteria and treatment recommendations.
  • Reviewed by a panel of eleven MDS experts using standardized scoring methods.
  • The best model, GPT-o3, had a 58% agreement rate with expert assessments.
  • Expert scores averaged up to 3.68 for GPT-o3 in diagnosis, indicating moderate performance.
  • Hallucinations were common, with error rates exceeding 25% across models.

Cite This Study

Consagra et al. (2025) studied this question.

synapsesocial.com/papers/69362f6e4fa91c937236e136https://doi.org/10.1182/blood-2025-4349
View Full Paper
Ask AI
Bookmark
Share