PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 10, 2026Journal of the American Medical Informatics Association1 citationsOpen Access

The detectability paradox: bilingual medical report generation with open-weight models and the limits of human oversight

View Full Paper
HRHossein RouhizadehASAbiram SandralegarAYAnthony Yazdani

Key Points

  • This research evaluates the quality of medical reports generated by large language models in English and French and examines human ability to distinguish between human-written and machine-generated reports.
  • Evaluated 4212 bilingual medical reports using text similarity metrics across multiple specialties.
  • Scoring conducted by a bilingual panel of certified physicians and medical residents, using a 1-5 Likert scale.
  • Performed a Turing-like test to assess human capability in identifying report origin.
  • Phi-4 model achieved the highest performance with ROUGE-1 score of 0.70 and BERTScore of 0.83.
  • Expert panel rated overall report quality at 4.6 out of 5, demonstrating high accuracy and fluency.
  • Automatic classifiers had an accuracy of 0.98, while human evaluators achieved only 0.60 in distinguishing report sources.

Abstract

OBJECTIVES: The automation of medical report generation using large language models (LLMs) could significantly reduce physicians' documentation burden while enhancing healthcare efficiency. However, the misuse of generative artificial intelligence in medical reporting can lead to important safety risks for patients. We addressed 2 questions: (1) What is the quality of medical reports generated by LLMs in English and French? and (2) Can we distinguish between human-written and LLM-generated medical reports? MATERIALS AND METHODS: We evaluated the quality of reports generated by several multilingual, open-weight LLMs using text similarity metrics on 4212 medical reports in English and French across multiple specialties. A bilingual expert panel of certified physicians (n = 4) and medical residents (n = 5) scored accuracy, fluency, and completeness of generated reports using a 1-5 Likert scale. Experts also completed a Turing-like test, blindly identifying reports as human or machine-generated. RESULTS: Phi-4 achieved the best overall performance (ROUGE-1: 0. 70, BERTScore: 0. 83). Expert evaluation confirmed high-quality reports in both languages (overall 4. 6/5. 0). Medical experts performed better than chance but struggled to differentiate human versus machine reports (accuracy: 0. 60). Automatic classifiers showed strong performance (accuracy: 0. 98). DISCUSSION: The high quality of LLM-generated reports supports their potential to enhance healthcare efficiency in multilingual settings. However, the discrepancy between human detection difficulty and automated detection success reveals inherent limitations in relying solely on human oversight for quality assurance and misuse prevention. CONCLUSIONS: Deployment of LLMs for medical reporting requires combining automated detection tools with human expertise to ensure patient safety. Dataset and code: https: //github. com/ds4dh/medicalᵣeportgeneration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rouhizadeh et al. (2026) studied this question.

synapsesocial.com/papers/6a0021fec8f74e3340f9cf5ehttps://doi.org/10.1093/jamia/ocag070
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1MIMIC-IV, a freely accessible electronic health record dataset2023 · 3,236 citations
  2. 2A Dataset for Evaluating Contextualized Representation of Biomedical Concepts in Language Models2024 · 9 citations
  3. 3Evaluating the effectiveness of biomedical fine-tuning for large language models on clinical tasks2025 · 37 citations
  4. 4The impact of the General Data Protection Regulation on health research2018 · 111 citations
  5. 5Reimagining Clinical Documentation With Artificial Intelligence2018 · 105 citations