Systematic review evaluates clinical utility and implementation challenges of large language models in healthcare, indicating important barriers to adoption.
Key Points
This review aims to assess the effectiveness and limitations of large language models in clinical healthcare, focusing on their performance and deployment challenges.
Systematic review of studies from January 2020 to January 2025
Eligibility for studies included transformer-based large language models with ≥10M parameters in clinical settings
Quality assessment using QUADAS-2, RE-AIM, and TRIPOD frameworks, with narrative synthesis per SWiM guidelines
Domain-adapted models achieved 88–98% accuracy on narrow tasks compared to 78–91% for general-purpose models
Real-world performance declined by 5–28% across clinical environments
Hallucination rates were 5–12% for domain-adapted models and 15–30% for general-purpose models in generative tasks