PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 14, 2026Open Forum Infectious Diseases0 citationsOpen Access

P-1976. Comparing Clinical Expertise and Chat-GPT in the Management of Septic Shock and Severe Pneumonia: A Pilot Study

View Full Paper
RBRhea BohraJSJassimran SinghESEric S. Silverman

Key Points

  • This pilot study aimed to compare the performance of ChatGPT-4 and physicians in managing septic shock and severe pneumonia.
  • Retrospective analysis of 50 cases at Saint Vincent Hospital
  • Comparison of physician and ChatGPT-4 outputs for investigations and antibiotics
  • Used Infectious Diseases Society of America guidelines as a standard
  • Utilized paired t-tests and McNemar’s test for evaluation
  • Statistical analysis conducted with Statistical Analysis Software version 9.4.
  • Physicians identified more appropriate investigations and antibiotics in severe pneumonia and septic shock
  • Mean difference of 1.45 tests and 0.91 antibiotics per case favored physicians for pneumonia
  • In septic shock, mean difference of 1 test and 1 antibiotic per case favored physicians
  • Similar MRSA coverage observed between both groups; physicians had slightly better accuracy
  • ChatGPT-4 showed potential in antimicrobial selection for Multi-Drug Resistant organisms.

Abstract

Abstract Background The evolution of artificial intelligence (AI) and large language models (LLMs) offers promising opportunities in infection management. Sepsis identification using AI has been integrated into many Electronic Medical Recording systems and applications for diagnostics and antimicrobial stewardship are emerging. This pilot study assessed ChatGPT-4® as a clinical decision aid for the management of septic shock and severe pneumonia.Flowchart showing methodology of pilot studyTable showing comparison of Physician and Chat-GPT4 performance Methods A retrospective study was conducted at Saint Vincent Hospital, Worcester on 50 cases (2023-2024). Physician-documented investigations and antibiotics were compared with ChatGPT-4® outputs generated using standard prompts with full blinding. Infectious Diseases Society of America guidelines served as the standard. Paired t-tests and McNemar’s test evaluated investigation appropriateness and antibiotic selection accuracy respectively, using Statistical Analysis Software® Version 9.4.Fig 3:Clinical template used to input H concordance was noted for Multi-Drug Resistant (MDR) organisms. Conclusion Physicians outperformed ChatGPT-4 across both conditions. However, ChatGPT-4® demonstrated comparable pathogen-specific antimicrobial selection and accuracy in MDR coverage, suggesting potential with diagnosis-specific prompting and antimicrobial stewardship. The results highlight the enduring importance of physician-led decision-making in an era increasingly shaped by AI. LLMs still face limitations including data privacy concerns and need for individualized contextual judgement that prevent their autonomous use. However, with refinement and responsible implementation, LLMs may evolve into a trusted aid to enhance physician decision-making, especially in areas with limited access to specialist care. This study has led to a prospective trial exploring ChatGPT-4’s use with real-time, targeted prompts throughout hospitalization. Disclosures All Authors: No reported disclosures

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bohra et al. (2026) studied this question.

synapsesocial.com/papers/6966f30613bf7a6f02c00871https://doi.org/10.1093/ofid/ofaf695.2143
Ask AI
Helpful
Bookmark
Share
View Full Paper