PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 9, 202511 citations

Uncertainty-aware Medical Diagnostic Phrase Identification and Grounding.

View Full Paper
KZKe ZouBYBai YangZhejiang Cancer HospitalBLBo LiuNingbo University

Key Points

  • Main finding: uMedGround represents a significant advancement in medical phrase grounding, addressing efficiency and trust issues.
  • Key evidence: Experimental results show uMedGround outperforms existing methods, enhancing diagnostic phrase detection and grounding accuracy.
  • Approach: The model employs a multimodal large language model with an embedded token for improved predictions and robustness.
  • Significance: This pioneering exploration into Medical Report Grounding promises better tools for clinicians, aiding accurate image analysis and interpretation.

Abstract

Medical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG) -which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, BOX, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zou et al. (2025) studied this question.

synapsesocial.com/papers/689dfe97d61984b91e13bfachttps://doi.org/10.1109/tpami.2025.3596878
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Heterogeneous Knowledge Grounding for Medical Question Answering with Retrieval Augmented Large Language Model2024 · 9 citations
  2. 2Interactive computer-aided diagnosis on medical image using large language models2024 · 144 citations
  3. 3Measurement Guidance in Diffusion Models: Insight from Medical Image Synthesis2024 · 30 citations
  4. 4Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment2023 · 21 citations
  5. 5ChestX-Ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases2017 · 3,481 citations