PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026Complex & Intelligent Systems0 citationsOpen Access

GLFormer: a hierarchical cross-layer transformer decoding framework with global–local feature collaboration for video action recognition

HWHanbo WuXMXin MaXLXiang Li

Key Points

  • This research aims to improve video action recognition by addressing limitations in global temporal and local motion dynamics modeling.
  • Developed a novel Transformer-based framework named GLFormer.
  • Utilized a frozen CLIP-pretrained Vision Transformer (ViT) for extracting spatial features from video frames.
  • Implemented hierarchical cross-layer decoding to process global and local information separately.
  • Applied temporal encoding on CLS tokens for global semantics and refined patch tokens for local motion cues.
  • GLFormer achieves superior performance on benchmark datasets compared to prior methods.
  • Significantly reduces computational overhead while maintaining high accuracy.
  • Demonstrates effective integration of global and local feature collaboration.

Abstract

Abstract CLIP’s effectiveness across image recognition tasks stems from its large-scale multimodal pretraining, which equips it with powerful and transferable visual representations. However, when directly applied to video action recognition, the CLIP image encoder often falls short in modeling global temporal dependencies and fine-grained local motion dynamics across frames. Moreover, by treating all patch tokens uniformly, it lacks the ability to attend to motion-salient regions that are critical for recognizing human actions. To address these limitations, we propose GLFormer, a novel Transformer-based framework that integrates Global–Local collaborative features of the whole video with a hierarchical cross-layer transformer decoding strategy. Specifically, we employ a frozen CLIP-pretrained Vision Transformer (ViT) as the image encoder to extract multi-layer spatial features from video frames. We then decouple the modeling of global and local information for each-layer features: CLS tokens are temporally encoded to capture global video-level semantics, while patch tokens are refined via spatial attention and temporal residual modeling to highlight local motion cues. The resulting spatio-temporal features from multiple ViT layers are progressively aggregated through a hierarchical Transformer decoder, where a learnable video-level query token interacts with features across different layers in a cross-layer manner. Extensive experiments on benchmark datasets demonstrate that GLFormer achieves remarkable performance while significantly reducing computational overhead, striking a strong balance between accuracy and efficiency for video action recognition.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wu et al. (2026) studied this question.

synapsesocial.com/papers/69e3201440886becb653f24ehttps://doi.org/10.1007/s40747-026-02314-3
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1HMDB: A large video database for human motion recognition2011 · 3,970 citations
  2. 2Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting2023 · 106 citations