Understanding the complexity of the rhythm and stylistic variability in how people actually perform their movements requires a computer-generated model that is capable of capturing all of this information utilizing multiple representations. Currently available skeleton-based action recognition systems can provide accurate recognition for many simple movements or actions. However, skeleton-based systems do not work well for distinguishing between the very small or detailed movements, or the expressive aspects, which are part of dance movements. To address this limitation, we propose a Dance Action Feature-Clustering and Adaptive Transformer (DAFCA-T) framework, which combines the use of differentiated clustering of dance behaviour with graph-aware spatiotemporal attention techniques, multi-level adaptation, and an innovative learning mechanism to obtain representations of dance movement patterns. We begin by encoding motion behaviours using a transformer network that takes into account the structure of the joints, which allows us to develop robust models of the relationships between joints and how they interact over time. We will then use a differentiable prototype clustering module to identify the collection of primitive movements that can be used to represent a dance. These prototype movements will provide an interpretable framework for the identification of dance actions. We also include a style adaptation module, which will combine FiLM modulated positional encodings (based on styles) with tempo-aware positional encodings and a lightweight test-time adaptive approach to refine internal representations without the need for annotated styles. DAFCA-T significantly outperforms the previously existing state-of-the-art methods for action recognition, key-frame detection and movement quality assessment based on extensive testing on the DanceMotion-3D, AIST Dance Motion Capture and NTU-RGBD datasets. The supplemental analyses also indicate that both the prototype learning and the adaptive modules provide complementary information for understanding dance action and demonstrate how the proposed prototype relates to the smallest motion patterns that comprise a dance. Consequently, the proposed approach is effective and interpretable for future research and has potentially broad applicability in the fields of sports science, advanced computational dance analysis and the analysis of how humans interact with computers.
No takes yet. Share an insight, caveat, or question.
Ning li (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: