Comparative survey reveals critical gaps in large language model factuality and audience adaptation for clinical trials, highlighting the need for specialized faithfulness guardrails.
Large language models (LLMs) have made automatic summarization of clinical-trial evidence practical, but a trustworthy trial summary must satisfy two demands that the literature usually studies apart: it must reach its intended audience - plain enough for a patient, precise enough for a clinician - and it must be factually faithful, since a fluent summary that misstates a dose or an outcome is a safety risk. This survey reviews twenty representative studies (2020-2026) across three threads: clinical-trial and medical summarization, audience-adaptive (plain-language) summarization, and claim factuality and hallucination evaluation. We compare them along six shared dimensions - model, dataset, audience, evaluation metric, headline result, and reported limitation. We find a consistent pattern: surface-level metrics fail in both directions, with overlap scores such as ROUGE not tracking faithfulness and readability formulas not tracking genuine understanding. Two concrete findings anchor this: even GPT-4 with chain-of-thought reaches only about 75% balanced accuracy in judging factual consistency, and no study in our set jointly targets audience-appropriateness and claim-level factuality for trial summarization. We name this three-way gap and outline the benchmarks and guardrails needed to close it, including clinically validated faithfulness metrics, comprehension-based evaluation, joint audience-and-factuality datasets, and audience-conditioned generation with a factuality guardrail. The survey offers a unified map of a fragmented field and a concrete agenda for trustworthy clinical-trial summarization.
No takes yet. Share an insight, caveat, or question.
Muhammad Zaheer Khan (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: