PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 20, 2025Journal of Imaging1 citationsOpen Access

Empirical Evaluation of Invariances in Deep Vision Models

View Full Paper
KKKonstantinos KeremisΕVΕleni VrochidouGPGeorge A. Papakostas

Key Points

  • Results demonstrate that vision transformers outperform convolutional neural networks under blur and noise, indicating model-specific strengths.
  • Segmentation models like SegFormer and Mask2Former show higher resilience to geometric variations, which challenges existing assumptions.
  • The study assesses robustness across three tasks: object localization, recognition, and semantic segmentation, utilizing key metrics.
  • Findings provide insights into real-world applicability, as both model types reveal vulnerabilities to rotation and extreme scale changes.

Abstract

The ability of deep learning models to maintain consistent performance under image transformations-termed invariances, is critical for reliable deployment across diverse computer vision applications. This study presents a comprehensive empirical evaluation of modern convolutional neural networks (CNNs) and vision transformers (ViTs) concerning four fundamental types of image invariances: blur, noise, rotation, and scale. We analyze a curated selection of thirty models across three common vision tasks, object localization, recognition, and semantic segmentation, using benchmark datasets including COCO, ImageNet, and a custom segmentation dataset. Our experimental protocol introduces controlled perturbations to test model robustness and employs task-specific metrics such as mean Intersection over Union (mIoU), and classification accuracy (Acc) to quantify models’ performance degradation. Results indicate that while ViTs generally outperform CNNs under blur and noise corruption in recognition tasks, both model families exhibit significant vulnerabilities to rotation and extreme scale transformations. Notably, segmentation models demonstrate higher resilience to geometric variations, with SegFormer and Mask2Former emerging as the most robust architectures. These findings challenge prevailing assumptions regarding model robustness and provide actionable insights for designing vision systems capable of withstanding real-world input variability.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Keremis et al. (2025) studied this question.

synapsesocial.com/papers/68d469c131b076d99fa662ddhttps://doi.org/10.3390/jimaging11090322
Ask AI
Helpful
Bookmark
Share
View Full Paper