Non-intrusive sign language recognition has started to redefine the wellness of millions of hearing impaired people around the world. In this study we propose two models, named ContiNet and F-ContiNet. ContiNet is a 3D convolutional network with a constant number of temporal snapshots, and F-ContiNet is an evolution of ContiNet where the fusion idea is applied so that the spatial and temporal information from the early layers will have greater direct impact on the final recognition. We have conducted extensive experiments on a 500-class dataset and its 100-class subset, using both raw videos and mask videos produced by the skeletal information obtained from OpenPose. For all the models we have tested, those using mask videos performed significantly better and learned faster than those using raw videos, and F-ContiNet performed the best, with the best per-class top-1 accuracy of 96.6% over the 500-class dataset. More importantly, it confirms that keeping a constant temporal dimension in a deep network is a viable approach to SLR, especially when the temporal information from all the available snapshots can be fully exploited by a time-sensitive construct such as ConvLSTM.
No takes yet. Share an insight, caveat, or question.
Han et al. (2021) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: