The transition to sixth-generation (6G) networks necessitates shifting from bit-level transmission to semantic communication for latency-sensitive Voice over Internet Protocol (VoIP) services, ensuring reliable human-centric interactions. However, conventional architectures process packets uniformly, causing severe perceptual degradation during network congestion, which severely limits overall user experience. This study develops a dynamic semantic prioritisation and real-time voice reconstruction framework optimised for 6G VoIP, addressing critical latency constraints. The research integrated a deep neural semantic encoder, a reinforcement learning scheduler, and a neural vocoder, evaluating performance via the LibriSpeech dataset in a 3GPP-compliant NS-3 simulation environment to ensure rigorous statistical validation. Quantitatively, the framework improved Perceptual Evaluation of Speech Quality by 34.2%, reduced latency by 28.5% to 42ms, and increased Short-Time Objective Intelligibility by 0.18 under 15% packet loss, demonstrating superior network resilience. Qualitatively, spectral analysis confirmed robust harmonic preservation and effective gap concealment, enabling critical phonetic data to bypass congestion bottlenecks efficiently, thereby maintaining natural speech prosody. Ultimately, integrating semantic extraction with cross-layer scheduling establishes a scalable architecture, advancing meaning-driven 6G communications. This innovation significantly enhances capacity without requiring additional spectrum allocation for operators, consequently paving the way for intelligent mobile voice services within the rapidly evolving global telecommunications network ecosystem today.
Solomon Malcolm Ekolama (Fri,) studied this question.