Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
March 21, 2026Tatra Mountains Mathematical PublicationsOpen Access

Contextualized Vision Transformers (CVT): Adaptive Spectral Embedding and Feature Gating for Precise Text-Graphics Classification

View Full Paper
Ask AI
Bookmark
Share

Authors

MGMridul GhoshKDKonrad DürrbeckRFRoland Fischer

Discussion

Loading...

Member takes

Overview

Innovative architecture enhances text-graphics classification, indicating improved processing efficiency.

Key Points

  • To develop an effective model for categorizing images into graphics-only, text-only, and mixed-content categories.
  • Introduced Contextualized Vision Transformer (CVT) architecture.
  • Employed Learnable Patch Decomposition (LPD) for efficient patch embedding extraction.
  • Implemented Adaptive Spectral Embedding (ASE) for dynamic spatial representation.
  • Integrated Contextual Feature Gating (CFG) for selective feature enhancement.
  • Utilized K-fold cross-validation to assess model robustness.
  • Achieved significant improvements in classification accuracy, precision, and recall.
  • Demonstrated superior performance over state-of-the-art Vision Transformers.
  • Highlighted effectiveness in content triage for large-scale visual processing.

Cite This Study

Ghosh et al. (2026) studied this question.

synapsesocial.com/papers/69be35606e48c4981c67398fhttps://doi.org/10.2478/tmmp-2026-0003
View Full Paper
Ask AI
Bookmark
Share