Recognizing subtle and complex dance movements requires models that can capture detailed spatial cues in individual frames while also tracking long-range temporal dynamics across entire performances. We propose a hybrid deep learning framework that first uses a modern convolutional neural network, pre-trained on large-scale image data, to extract multi-scale spatial features from each video frame and then applies a bidirectional recurrent neural network with a temporal attention mechanism to emphasize the most informative motion segments when aggregating features over time. The model is trained and evaluated on a video dataset of traditional Balinese dance movements, where it achieves higher classification accuracy and greater robustness to viewpoint changes and partial occlusion than several existing deep learning and classical machine learning baselines. Ablation experiments show that both the strong spatial feature extractor and the temporal attention module contribute significantly to performance gains. The resulting framework is compact, data-efficient, and easily adaptable to other cultural dance archives and more general tasks involving spatiotemporal human action recognition.
Zhang et al. (Tue,) studied this question.