Deep learning-based sign language recognition plays a pivotal role in facilitating communication for the deaf community. Current approaches, while effective, often introduce redundant information and incur excessive computational overhead through global feature interactions. To address these limitations, this paper introduces a Deformable Correlation Network (DCA) designed for efficient temporal modeling in continuous sign language recognition. The DCA integrates a Deformable Correlation (DC) module that leverages spatio-temporal driven offsets to adjust the sampling range adaptively, thereby minimizing interference. Additionally, a multi-scale local sampling strategy, guided by motion prior, enhances temporal modeling capability while reducing computational costs. Furthermore, an attention-based Correlation Matrix Filter (CMF) is proposed to suppress interference elements by accounting for feature motion patterns. A long-term temporal enhancement module, based on spatial aggregation, efficiently leverages global temporal information to model the performer’s holistic limb motion trajectories. Extensive experiments on three benchmark datasets demonstrate significant performance improvements, with a reduction in Word Error Rate (WER) of up to 7.0% on the CE-CSL dataset, showcasing the superiority and competitive advantage of the proposed DCA algorithm.
No takes yet. Share an insight, caveat, or question.
Jiang et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: