Current sEMG-based speech generation methods primarily depend on extensive datasets from individual participants, which presents a challenge for users. Furthermore, prior research methodologies frequently necessitate synchronous recordings of sEMG speech for model training, rendering them inappropriate for patients with speech impairments. This article presents a cross-subject sEMG-to-speech (ETS) conversion system utilising content features and model calibration techniques. This technique employs a pre-trained acoustic model to derive speaker-independent acoustic features and modifies the model to accommodate the EMG characteristics of new subjects via Child Tune model calibration. To create an ETS system that works for people who have trouble with language, we suggest using electronic synthesised audio to train the ETS model instead of a real person's voice and then using a voice encoder with speech conversion to put the speech back together. The experimental results indicate that our proposed sEMG-to-speech (ETS) conversion system can attain a Character Error Rate (CER) of 21.71% following calibration with 20 minutes of data from new users. Furthermore, utilising electronically synthesised audio for training the ETS model yields a CER comparable to that obtained with human voice.
Khanna et al. (Wed,) studied this question.