This research offers a brand-new framework for producing high-quality, photo-realistic sign language videos using sign language corpora. The framework consists of three main components such as Neural Machine Translation (NMT) using transformer, an open pose estimation approach, and a Generative Adversarial Network (GAN) with VGG-19 classifier. The NMT transformer translates the sign language corpus into a sequence of sign glosses, pose estimation model performs mapping of equivalent signer images and further assists to generate sign images. The GAN with VGG-19 classifier then generates high- quality sign language videos from the sign language images. The proposed method performs at the cutting edge in terms of evaluation on three benchmark dataset using metrics such as PSNR, SSIM, FID, and Inception Score. In addition, the framework offers better translation outcomes with less processing overhead and produces high-quality, photo-realistic sign language videos without the need for animation or avatar techniques. The proposed framework has the potential to offer the uninterrupted communication medium between common people with deaf-mute society. The evaluation scores highlights the improved perfor- mance of the proposed model compared with earlier approaches in terms of visual quality.
No takes yet. Share an insight, caveat, or question.
R et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: