PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Beyond Black-Box AI: A Quantitative Grad-CAM Analysis of Convolutional Neural Network Interpretability in COVID-19 Chest X-Ray Classification

View Full Paper
ASAiman Abd SaeedRRRasber Dhahir RashidSalahaddin University-Erbil

Key Points

  • This research aims to quantify the interpretability of CNNs using Grad-CAM in COVID-19 chest X-ray classification.
  • Evaluated six CNN models: VGG16, VGG19, ResNet-101, NASNet-Mobile, NASNet-Large, and Xception.
  • Implemented an automated pipeline for generating objective metrics from Grad-CAM heatmaps and lung masks.
  • Classifications measured using accuracy, precision, recall, and F1-score, alongside IoU and Dice scores for model interpretability.
  • Xception achieved the highest accuracy at 95.90% with an F1-score of 95.92%.
  • VGG19 reached the highest precision of 98.89%.
  • Classification accuracies across models ranged from 90% to 96% with notable anatomical interpretability scores.

Abstract

Modern AI models use deep architectures that obscure how predictions are made. Without understanding how models reach their predictions, it becomes difficult to verify reasoning, identify biases, or trust their reliability in high-stakes domains like healthcare. Many COVID-19 chest X-ray (CXR) studies report high accuracy and present qualitative gradient-weighted class activation mapping (Grad-CAM) heatmaps, providing no quantitative evidence of alignment with lung anatomy and relying on manual, subjective inspection. We introduce an automated quantitative pipeline that converts interpretability into objective, anatomy grounded metrics between Grad-CAM heatmaps and lung masks. We evaluate six convolutional neural networks (CNNs): VGG16, VGG19, ResNet-101, NASNet-Mobile, NASNet-Large, and Xception, for both classification performance and anatomical interpretability in COVID-19 CXR detection. Classification accuracies ranged from 90% to 96%, with Xception achieving the highest accuracy (95.90%) and a balanced precision, recall, and F1-score of 95.92%. NASNet-Large and VGG19 followed at 94.87%, with VGG19 reaching the highest precision (98.89%). To assess model transparency, we automated interpretability analysis by thresholding the Grad-CAM outputs and comparing them to radiologist-annotated lung masks using Intersection-over-Union (IoU) and Dice score metrics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Saeed et al. (2026) studied this question.

synapsesocial.com/papers/69f6e5cf8071d4f1bdfc66d2https://doi.org/10.21271/zjpas.38.2.13
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Exploration of Interpretability Techniques for Deep COVID-19 Classification Using Chest X-ray Images2024 · 30 citations
  2. 2Machine-Learning-Enabled Diagnostics with Improved Visualization of Disease Lesions in Chest X-ray Images2024 · 10 citations
  3. 3Enhancing Reliability in Deep Learning Diagnosis of Brain Tumors Using Grad-CAM2026
  4. 4Comparative Analysis of Visual Explainable AI Techniques for Chest X-ray Classification2026
  5. 5Hybrid ConvNeXtV2–ViT Architecture with Ontology-Driven Explainability and Out-of-Distribution Awareness for Transparent Chest X-Ray Diagnosis2026 · 3 citations