PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 13, 2026SHILAP Revista de lepidopterología3 citationsOpen Access

Tests of large language models' medical competence and application for clinical decision support of musculoskeletal rehabilitation

View Full Paper
RLRuikang LiuQLQiaoling LiuYLYi Li

Key Points

  • This study investigates the effectiveness of large language models in clinical applications related to musculoskeletal rehabilitation.
  • Involved 8 primary doctors and therapists testing 10 large language models (LLMs) in the first test.
  • 5 senior doctors and therapists evaluated responses in the second test.
  • 5 primary therapists acted as examinees in the third test.
  • Assessed quality of case analysis based on six dimensions including clinical reasoning and treatment plan accuracy.
  • Only ERNIE Bot X1 Turbo and Doubao 1.5 pro showed over 90% accuracy in the first test.
  • Chinese LLMs had fewer incorrect answers compared to English LLMs (9.6% vs. 14.8%).
  • Doubao 1.5 pro scored high in case understanding and clinical reasoning.
  • Primary therapists achieved a mean accuracy of 76.9% in the third test, with Doubao 1.5 pro reaching 85.8%.

Abstract

Objective Large language models (LLMs) are currently abundant and diverse, yet clinicians lack clarity on top performers, with uncertainty about general LLMs' expertise in musculoskeletal rehabilitation. This study aims to investigate the potential and correctness of LLMs in clinical application, and to evaluate whether LLMs could assist primary rehabilitation therapists to prepare for rehabilitation examination. Method 8 primary doctors and therapists tested 10 LLMs in the first test, 5 senior doctors and therapists assessed answers in the second test, and 5 primary therapists acted as examinees in the third test. We assessed the quality of case analysis based on six different dimensions, including Case Understanding, Clinical Reasoning, Primary Diagnosis, Differential Diagnosis, Treatment Plan Accuracy and Safety, and Guidelines Consensus. Results In the first test, only ERNIE Bot X1 Turbo and Doubao 1.5 pro had accuracy rates of over 90%, and Chinese LLMs had significantly fewer incorrect questions than English LLMs (9.6% vs. 14.8%, P 0.001). In the second test, Doubao 1.5 pro achieved relatively high scores in both cases, and LLMs gained high scores in “Case understanding”, “Clinical Reasoning” and “Diagnosis”. In the third test, primary therapists achieving a mean accuracy rate of 76.9%, and Doubao 1.5 pro improved its accuracy rates to 85.8%. Conclusions Doubao 1.5 pro possessed competent ability and application prospects, and was assessed as the best LLM for answering musculoskeletal rehabilitation questions. We also demonstrated that the response quality of local-language LLMs was significantly better than that of English LLMs in answering localized language questions.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/698ebeb185a1ff6a93016120https://doi.org/10.3389/fdgth.2025.1719340
Ask AI
Helpful
Bookmark
Share
View Full Paper