PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026PLOS Digital Health2 citationsOpen Access

Visual recognition limitations in multimodal large language models: A comparative analysis of histological image interpretation

View Full Paper
VMVolodymyr MavrychEYEinas M. YousefAYAhmed Yaqinuddin

Key Points

  • The research aims to evaluate how well multimodal LLMs perform in interpreting histological images.
  • Assessment of four multimodal LLMs using 144 histological images
  • Images represented four tissue types across three magnification levels
  • Evaluation based on three standardized questions regarding tissue identification
  • Grading of responses by three expert faculty members using a 4-point scale
  • Statistical analyses including Friedman tests and inter-rater reliability assessments
  • Gemini showed the highest performance with a mean score of 3.35 out of 4.00
  • Copilot and GPT-4o ranked second with scores of 2.76 each
  • Claude had the lowest performance score of 2.55
  • Inter-model variation was highest in epithelial tissues
  • Good inter-rater reliability was confirmed with an ICC greater than 0.85

Abstract

Multimodal large language models (LLMs) with image recognition capabilities have emerged as potential tools for medical image analysis, yet their performance in specialized domains like histology remains largely unexplored. The objective of this study was to systematically evaluate the performance of leading multimodal LLMs in histological image interpretation and assess their visual recognition capabilities. Four multimodal LLMs (GPT-4o, Claude Sonnet 4, Gemini 2.5 Flash, and Copilot) were evaluated using 144 histological images representing four tissue types (epithelial, connective, muscle, and nervous) at three magnification levels. Each image was assessed using three standardized questions: tissue identification, morphological features, and functional analysis. Three expert faculty members independently graded responses using a 4-point scale (1 = Poor to 4 = Excellent). Friedman tests, ICC, and post-hoc power analyses were performed with statistical significance set at p 0.85), confirming assessment consistency. Post-hoc power analysis validated statistical significance for primary comparisons but indicated insufficient power to distinguish between the three lower-performing models. Current multimodal LLMs exhibit significant limitations in visual recognition relative to text processing performance. The substantial cross-modal performance gaps reveal some constraints in visual processing architectures, though the underlying mechanisms require further investigation. These findings establish technical benchmarks for multimodal LLM development and highlight the need for specialized visual processing innovations in their imaging processes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mavrych et al. (2026) studied this question.

synapsesocial.com/papers/69be35d76e48c4981c6743d6https://doi.org/10.1371/journal.pdig.0001306
Ask AI
Helpful
Bookmark
Share
View Full Paper