ABSTRACT Skeleton‐based hand gesture recognition (SGR) has emerged as a prominent research area in computer vision (CV) due to the unique advantages of skeletal data. Although gesture recognition (GR) has traditionally relied on video streams and RGB image data, skeleton‐based approaches remain underexplored. Despite advances in deep learning techniques, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), graph convolutional networks (GCNs), attention mechanisms, and multimodal fusion, a systematic and comprehensive review of these methods within the context of SGR is lacking. This paper addresses this gap by first underscoring the significance of GR and the pivotal role of skeletal data in its analysis. We then present the data acquisition techniques for SGR and examine its key applications, such as human‐robot interaction (HRI), sign language recognition, gaming and entertainment, and healthcare and rehabilitation. Furthermore, we analyze state‐of‐the‐art methodologies and compare their strengths and limitations in detail. Based on a thorough evaluation of existing research, we identify current challenges and unresolved bottlenecks. Finally, we propose future research directions to advance SGR, focusing on model robustness, multimodal fusion, and dataset expansion.
Li et al. (Thu,) studied this question.