In the real world, Human Activity Recognition (HAR) remains challenging due to issues such as occlusion, dynamic backgrounds, and visual noise. Traditional models, such as CNNs, RNNs, and ST-GCNs, have constraints, including a small receptive field and the use of local features, which also reduce generalisation. We present ViT-HAR, a Vision Transformer framework that learns global spatio-temporal interactions and proposes two new modules, namely Contextual Patch Reweighting (CPR) and Attention-Guided Occlusion Masking (AGOM) to solve this issue. These parts enable selective attention of motion-relevant and nonoccluded areas, which increases robustness and interpretability in cluttered scenarios. In contrast to the previous versions of Vision Transformer architecture (ex, TimeSformer and ViViT), which use fixed attention by default, ViT-HAR is based on adaptive attention, redistributing contextual patches and masking unseen areas with dynamically varying weights in attempts to retain semantically salient information. The combined pipeline utilises dynamic frame sampling, contextual reweighting, and occlusion-based masking, resulting in an optimal trade-off between spatial and temporal coherence. NTU RGB+D and Kinetics-700 and UCF101 evaluation results indicate a maximum 6.5% greater Top-1 and show better F1-scores than 3D-CNN and RNN hybrids. Visualization of attention maps attests to the fact that ViT-HAR pays attention to meaningful motion signals, which is why the algorithm proves useful in healthcare monitoring, smart surveillance, and AR/VR. Lightweight and multimodal extensions are investigated in future work.
Mewada et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: