PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 11, 2026Hong Kong Journal of Emergency Medicine2 citationsOpen Access

Automated systematic reviews using machine learning and large language models in clinical practice guideline development: A perspective

View Full Paper
TOTakehiko OamiYOYohei OkadaTNTaka‐aki Nakada

Key Points

  • This perspective focuses on the integration of machine learning and large language models in automating systematic reviews for clinical guideline development.
  • Described the application of machine learning and large language models for citation screening in guideline development.
  • Evaluated the performance of ML-assisted tools in reducing screening time across clinical questions.
  • Discussed the role of prompt engineering in optimizing model performance for sensitivity and specificity.
  • Conducted comparative evaluations of multiple large language models to assess trade-offs in performance.
  • Machine learning-based citation screening reduced screening time, though effectiveness varied by clinical question.
  • Large language model-assisted screening improved efficiency without requiring task-specific training data.
  • Evaluations revealed trade-offs between sensitivity and specificity, impacting model selection based on task priorities.
  • Emerging evidence supports the role of ML and LLMs in other stages like search strategy development and quality assessment.

Abstract

Abstract Background Systematic reviews (SRs) are essential for the development of clinical practice guidelines but require time and human resources, raising concerns about sustainability and timeliness. Recent advances in machine learning (ML) and large language models (LLMs) offer promising opportunities to automate SR tasks, including citation screening. However, optimal strategies for integrating these technologies into guideline development workflows remain unclear. Main body Based on experiences from the SR automation team for the Japanese Clinical Practice Guidelines for Management of Sepsis and Septic Shock 2024, this perspective describes the practical application of ML‐ and LLM‐assisted approaches in guideline development. Semi‐automated citation screening using ML‐based tools reduced screening time, although performance varied across clinical questions. More recently, LLM‐assisted citation screening demonstrated further efficiency gains and improved flexibility without requiring task‐specific training data. Prompt engineering played a critical role in optimizing model performance, enabling sensitivity to be increased while preserving specificity. Comparative evaluations across multiple LLMs revealed inherent trade‐offs between sensitivity and specificity, highlighting the importance of model selection based on task priorities. Beyond citation screening, emerging evidence supports the potential role of ML and LLMs in search strategy development, data extraction, and quality assessment, although full automation remains limited by interpretability, domain variability, and the need for expert oversight. Conclusions ML‐ and LLM‐assisted automation is expected to reduce workload and enhance the efficiency of SRs in clinical guideline development. A hybrid human–AI approach, combining automated processes with expert judgment, represents a practical and safe pathway toward sustainable, timely, and reproducible guideline development.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Oami et al. (2026) studied this question.

synapsesocial.com/papers/698c1bef267fb587c655df17https://doi.org/10.1002/hkj2.70085
Ask AI
Helpful
Bookmark
Share
View Full Paper