PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 30, 2025npj Digital Medicine11 citationsOpen Access

Understanding the robustness of vision-language models to medical image artefacts

View Full Paper

Key Points

  • Images with strong artefacts were detected at low rates, showing VLMs' limitations.
  • On original unaltered images, VLMs achieved moderate accuracy of 0.645 in brain MRI scans.
  • The analysis involved evaluating VLMs across multiple artefact categories from real-world medical datasets.
  • Findings underscore the need for artefact-aware design in VLM development, aiming for improved robustness.

Abstract

Abstract Vision-language models (VLMs) show promise for answering clinically relevant questions, but their robustness to medical image artefacts remains unclear. We evaluated VLMs’ robustness through their performance on images with and without weak artefacts across five artefact categories, as well as their ability to detect images with strong artefacts. We built evaluation benchmarks using brain MRI scans, Chest X-ray, and retinal images, involving four real-world medical datasets. VLMs achieved moderate accuracy on original unaltered images (0.645, 0.602 and 0.604 for MRI, OCT, and X-ray applications, respectively). Accuracy declined with weak artefacts (−3.34%, −9.06% and −10.46%), while strong artefacts were detected at low rates (0.194, 0.128 and 0.115). Our findings indicated that VLMs are not yet capable of performing tasks on medical images with artefacts, underscoring the need to establish uniform benchmark thoroughly examining model robustness to image artefacts, and explicitly incorporate artefact-aware method design and robustness tests into VLM development.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

A 2025 study studied this question.

synapsesocial.com/papers/692b9d931d383f2b2a379e2dhttps://doi.org/10.1038/s41746-025-02108-w
Ask AI
Helpful
Bookmark
Share
View Full Paper