PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 11, 2026JMIR Formative Research0 citationsOpen Access

Personalized Diabetes Treatment Support Using Large Language Models Fine-Tuned on Electronic Health Records: Development and Evaluation Study

View Full Paper
SHshenyang heYZYu ZhangJLJiaxi Li

Key Points

  • The study aims to develop and assess a large language model-based system for personalized diabetes treatment support.
  • Three compact large language models were fine-tuned on deidentified outpatient electronic health records.
  • Models were optimized using a parameter-efficient low-rank adaptation approach.
  • The optimized models were integrated into a hospital information system using a retrieval-augmented generation framework.
  • The GLM4-9B model showed the best performance in generating treatment plans and laboratory recommendations.
  • Achieved a mean score of 67.93 on the Bilingual Evaluation Understudy for 4-grams.
  • Lower mean scores were noted for Recall-Oriented Understudy evaluations for unigrams and bigrams.

Abstract

Abstract Background Effective diabetes management requires individualized treatment strategies tailored to patients’ clinical characteristics. With recent advances in artificial intelligence, large language models (LLMs) offer new opportunities to enhance clinical decision support, particularly in generating personalized recommendations. Objective This study aimed to develop and evaluate an LLM-based outpatient treatment support system for diabetes and examine its potential value in routine clinical decision-making. Methods Three compact LLMs (Llama 3.1-8B, Qwen3-8B, and GLM4-9B) were fine-tuned on deidentified outpatient electronic health records using a parameter-efficient low-rank adaptation approach. The optimized models were embedded into a prototype hospital information system via a retrieval-augmented generation framework to generate individualized treatment recommendations, laboratory test suggestions, and medication prompts based on demographic and clinical data. Results Among the models evaluated, the fine-tuned GLM4-9B demonstrated the strongest performance, producing clinically reasonable treatment plans and appropriate laboratory test recommendations and medication suggestions. It achieved a mean Bilingual Evaluation Understudy for 4-grams score of 67.93 (SD 2.74) and mean scores of 44.30 (SD 3.91) for Recall-Oriented Understudy for Gisting Evaluation for overlap of unigrams, 27.34 (SD 1.85) for Recall-Oriented Understudy for Gisting Evaluation for overlap of bigrams, and 37.67 (SD 2.88) for Recall-Oriented Understudy for Gisting Evaluation for Longest Common Subsequence. Conclusions The fine-tuned GLM4-9B shows strong potential as a clinical decision support tool for personalized diabetes care. It can provide reference recommendations that may improve clinician efficiency and support decision quality. Future work should focus on enhancing medication guidance, expanding data sources, and improving adaptability in cases involving complex comorbidities.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

he et al. (2026) studied this question.

synapsesocial.com/papers/698c1cd3267fb587c655f890https://doi.org/10.2196/71541
Ask AI
Helpful
Bookmark
Share
View Full Paper