OpenAI o3 demonstrated the highest accuracy (87.6%) among nine generative AI models for perioperative medication discontinuation decisions, followed by Gemini 2.5 Pro Exp (84.8%) and GPT-4o (83.8%).
How accurate are various generative AI models in determining perioperative medication discontinuation?
先进的生成性人工智能模型可以在约85%的情况下准确支持围手术期药物停用决策,但由于10-20%的错误率仍需药剂师验证。
Generative artificial intelligence (AI) has rapidly advanced and is expected to enhance the efficiency and quality of clinical practice. In pharmacy practice, determining perioperative medication discontinuation is a critical task that is directly related to patient safety. However, standardized guidance remains limited, particularly for over-the-counter (OTC) drugs and dietary supplements, and decisions often rely on the individual expertise of pharmacists. In this study, the accuracy of nine generative AI models available in April 2025 (GPT-4o, GPT-4o mini, OpenAI o3, Gemini 2.5 Pro Exp, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Llama 4 Scout, and DeepSeek R1) for perioperative medication discontinuation decisions was evaluated using 15 mock prescription sets comprising 105 items. Each model received a standardized Japanese prompt, and outputs were independently assessed by five hospital pharmacists (≥5 years of clinical experience) based on three criteria: accurate drug identification, appropriateness of discontinuation and resumption timing, and the validity of the pharmacological rationale. A response was considered correct when at least four of the pharmacists agreed. OpenAI o3 demonstrated the highest accuracy (87.6%), followed by Gemini 2.5 Pro Exp (84.8%) and GPT-4o (83.8%). Lightweight models demonstrated lower accuracy, particularly for OTC products, dietary supplements, and fixed-dose combination drugs. High-performance models with advanced reasoning capabilities exhibited high accuracy and may serve as useful decision-support tools. However, incorrect responses occurred in approximately 10 – 20% of cases, even among the top-performing models. Therefore, safe clinical implementation requires careful model selection, integration with institutional knowledge resources, and final verification by pharmacists.
Makieda等(周四)在围手术期药物停用中进行了其他研究(n=105)。生成性人工智能模型(OpenAI o3、Gemini 2.5 Pro Exp、GPT-4o等)与临床药剂师共识(参考标准)在围手术期药物停用决策的准确性上进行了评估。OpenAI o3在九种生成性人工智能模型中显示出最高准确性(87.6%),其次是Gemini 2.5 Pro Exp(84.8%)和GPT-4o(83.8%)。