PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 17, 2026Sakarya University Journal of Computer and Information Sciences0 citationsOpen Access

ViT-HAR: Vision Transformer-Based Human Activity Recognition in Cluttered Environments

AMArvind MewadaSAShahnawaz AhmadMAMohd Aquib Ansari

Key Points

  • The aim is to improve human activity recognition in challenging, cluttered environments using a vision transformer framework.
  • Developed the ViT-HAR framework incorporating contextual patch reweighting and attention-guided occlusion masking.
  • Utilized dynamic frame sampling for adaptive attention in visual data processing.
  • Evaluated performance on NTU RGB+D, Kinetics-700, and UCF101 datasets.
  • Achieved a maximum improvement of 6.5% in Top-1 accuracy.
  • Outperformed traditional 3D-CNN and RNN hybrids in F1-scores.
  • Visualization of attention maps confirmed the ability to focus on relevant motion signals.

Abstract

In the real world, Human Activity Recognition (HAR) remains challenging due to issues such as occlusion, dynamic backgrounds, and visual noise. Traditional models, such as CNNs, RNNs, and ST-GCNs, have constraints, including a small receptive field and the use of local features, which also reduce generalisation. We present ViT-HAR, a Vision Transformer framework that learns global spatio-temporal interactions and proposes two new modules, namely Contextual Patch Reweighting (CPR) and Attention-Guided Occlusion Masking (AGOM) to solve this issue. These parts enable selective attention of motion-relevant and nonoccluded areas, which increases robustness and interpretability in cluttered scenarios. In contrast to the previous versions of Vision Transformer architecture (ex, TimeSformer and ViViT), which use fixed attention by default, ViT-HAR is based on adaptive attention, redistributing contextual patches and masking unseen areas with dynamically varying weights in attempts to retain semantically salient information. The combined pipeline utilises dynamic frame sampling, contextual reweighting, and occlusion-based masking, resulting in an optimal trade-off between spatial and temporal coherence. NTU RGB+D and Kinetics-700 and UCF101 evaluation results indicate a maximum 6.5% greater Top-1 and show better F1-scores than 3D-CNN and RNN hybrids. Visualization of attention maps attests to the fact that ViT-HAR pays attention to meaningful motion signals, which is why the algorithm proves useful in healthcare monitoring, smart surveillance, and AR/VR. Lightweight and multimodal extensions are investigated in future work.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mewada et al. (2026) studied this question.

synapsesocial.com/papers/69b8f10fdeb47d591b8c5d4fhttps://doi.org/10.35377/saucis...1745614
Ask AI
Helpful
Bookmark
Share
View Full Paper