Sign language plays a vital role in enabling interaction for individuals with hearing and speech impairment, making accurate alphabet-level recognition a fundamental requirement for accessible human-computer interaction systems. This paper investigates visual deep learning approaches for interpreting sign language alphabets from image data. A systematic framework is developed to assess the effectiveness of transfer learning-based convolutional neural networks in capturing discriminative hand gesture features. Using a benchmark American Sign Language (ASL) alphabet dataset, several advanced architectures, including ConvNeXtXLarge, EfficientNet, VGG19, and ResNet-50, are examined under a unified experimental protocol. The comparative analysis reveals that ConvNeXtXLarge achieves the uppermost recognition performance, attaining an accuracy of 99.81%, while EfficientNet, VGG19, and ResNet-50 also demonstrate strong results with accuracy of 99.68%, 99.31%, and 97.29%, respectively. These findings emphasize the effectiveness of modern transfer learning strategies in enhancing visual representation learning for fine-grained motion recognition. The proposed evaluation framework offers practical insights into model selection for scalable and reliable sign language interpretation systems, contributing to the advancement of inclusive assistive technologies and real-world visual language understanding applications.
Nath et al. (Sat,) studied this question.