PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 2025Journal of Evaluation in Clinical Practice12 citations

Evaluating AI‐Generated Meal Plans for Simulated Diabetes Profiles: A Guideline‐Based Comparison of Three Language Models

View Full Paper
HBHati̇ce Merve BayramGelişim ÜniversitesiSASedat ArslanBursa Uludağ Üni̇versi̇tesi̇AÖArda ÖztürkcanGelişim Üniversitesi

Key Points

  • ChatGPT‐4.1 showed 70.9% alignment with energy requirements, but overestimated fat intake.
  • Grok‐3 achieved 83.1% energy accuracy while missing several micronutrient targets.
  • DeepSeek adjusted protein based on BMI, yet underdelivered carbohydrates in its meal plans.
  • Integration of retrieval-augmented generation may improve AI performance in dietary planning.

Abstract

ABSTRACT Aims This synthetic simulation, using no real patient data, study aimed to evaluate and compare the performance of three prominent large language models (LLMs)—ChatGPT‐4.1, Grok‐3 and DeepSeek—in generating medical nutrition therapy aligned dietary plans for adults with type 2 diabetes mellitus (T2DM). Methods A simulation‐based design was employed using 24 standardized virtual patient profiles differentiated by gender and body mass index (BMI) category. Each LLM was prompted in Turkish to generate 3‐day meal plans. Outputs were assessed for energy and macro‐/micronutrient accuracy, adherence to national and international T2DM guidelines and alignment with the nutrition care process (NCP). Results ChatGPT‐4.1 showed the highest alignment with energy requirements (70.9%) but overestimated fat intake. Grok‐3 demonstrated superior energy accuracy (83.1%) but failed to meet several micronutrient targets. DeepSeek adjusted protein intake according to BMI but underdelivered carbohydrates. None of the models demonstrated full concordance with the NCP framework, particularly in the diagnosis and monitoring components. Frequent hallucinations and lack of clinical contextualization were noted. Integration of retrieval‐augmented generation (RAG) was identified as a potential improvement strategy. Conclusion While LLMs showed promise in generating baseline dietary guidance in a simulated context, these results reflected concordance with guideline documents only and concordance with guideline documents only and should not be interpreted as evidence of equivalence to dietitian‐led care. These findings reflected model behaviour in synthetic scenarios only and highlighted the need for RAG integration and expert supervision before any clinical application.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bayram et al. (2025) studied this question.

synapsesocial.com/papers/68e80eb363e2e2f707877c0bhttps://doi.org/10.1111/jep.70295
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Large Language Models as Clinical Nutrition Decision Tools: Quantitative Bias and Guideline Deviation in Type 2 Diabetes Meal Planning2026 · 2 citations
  2. 2An Analysis of the Nutritional Value of Diets Generated by Large Language Models Under Average User Conditions—3(LM)Diet Study2026
  3. 3Assessment of large language model chatbots for hemodialysis meal planning: a descriptive study2026
  4. 4Evaluating the Effectiveness and Safety of Large Language Model in Generating Type 2 Diabetes Mellitus Management Plans: A Comparative Study with Medical Experts Based on Real Patient Records2024
  5. 5Evaluating Large Language Models-Generated Health Education Materials for Discharged Patients with Diabetes: A Comparative Analysis.2026