PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 9, 20250 citations

Plain Language Adaptations of Biomedical Text Using LLMs: Comparision of Evaluation Metrics.

View Full Paper
PKPrimož KocbekLKLeon KopitarGŠGregor Štiglic

Key Points

  • The core finding shows that gpt-4o-mini outperformed other approaches in simplifying biomedical texts for better health literacy.
  • Evaluation used quantitative metrics like the Flesch-Kincaid grade level and qualitative metrics via 5-point Likert scales.
  • The study developed different approaches: baseline, AI agent, and fine-tuning using public biomedical datasets.
  • Results suggest G-Eval is a promising LLM-based quantitative metric aligning closely with qualitative evaluations.

Abstract

This study investigated the application of Large Language Models (LLMs) for simplifying biomedical texts to enhance health literacy. Using a public dataset, which included plain language adaptations of biomedical abstracts, we developed and evaluated several approaches, specifically a baseline approach using a prompt template, a two AI agent approach, and a fine-tuning approach. We selected OpenAI gpt-4o and gpt-4o mini models as baselines for further research. We evaluated our approaches with quantitative metrics, such as Flesch-Kincaid grade level, SMOG Index, SARI, and BERTScore, G-Eval, as well as with qualitative metric, more precisely 5-point Likert scales for simplicity, accuracy, completeness, brevity. Results showed a superior performance of gpt-4o-mini and an underperformance of FT approaches. G-Eval, a LLM based quantitative metric, showed promising results, ranking the approaches similarly as the qualitative metric.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kocbek et al. (2025) studied this question.

synapsesocial.com/papers/689dfe97d61984b91e13befbhttps://doi.org/10.3233/shti250946
Ask AI
Helpful
Bookmark
Share
View Full Paper