Key points are not available for this paper at this time.
With the increasing complexity of next generation network applications and the coexistence of diverse service requirements, Generative AI (GAI) and Large Models (LMs) based semantic communication are widely regarded as promising solutions to address these challenges. The goal of these systems is not only to reduce system burden by reducing transmission data, but also to adapt to new and complex requirements. In this paper, we propose a semantic communication system designed to meet diverse requirements of speech applications while enabling accurate speech transmission. The semantic encoder comprises an unsupervised model wav2vec 2.0 for learning universal speech representations to enable adaptability across various speech-related tasks. It also includes a prosodic feature encoder from the style embedding module of Global Style Tokens (GST) Tacotron. The semantic decoder integrates a phoneme recognition module and a GST-Tacotron-based text-to-speech (TTS) module to facilitate accurate and expressive reconstruction of the original speech signal, with the incorporation of prosodic features enhancing the naturalness and intelligibility of the synthesized speech. The proposed system has been tested in noisy channels. It demonstrates that the system maintains superior and robust performance even at Bit Error Rate (BER) of 10−1, as reflected by a stable Character Error Rate (CER) approximately 0.0940 and 0.0649 for the base and large versions of wav2vec 2.0 respectively in speech recognition, and consistent cFDSD scores approximately 0.9 in speech quality assessment. This performance surpasses that of the existing semantic communication systems, while also providing reliable support for a wider range of downstream speech applications.
Wang et al. (Thu,) studied this question.