The compartmentalization of encryption protocol designs and evaluations for audio and video, rooted in the inherent differences in data dimensionality, undermines their universality and interoperability, thereby imposing substantial implementation overhead and complexity on integrated multimedia security solutions. To bridge this gap, we propose a cross-modal audio encryption protocol that converts linear audio samples into two-dimensional byte matrices via a straightforward modal transformation and encrypts them based on Shannon’s confusion–diffusion framework. Methodologies analogous to those employed in video encryption, together with conventional audio analysis techniques, are utilized to cross-validate its statistical properties and security. Experimental results demonstrate that the proposed protocol exhibits desirable statistical properties and provides robust resistance against various attacks. Furthermore, its deployment in a three-user real-time secure voice conferencing system over a campus network verifies its feasibility and practical applicability. This paper validates that video encryption algorithms, together with their sophisticated evaluation methodologies, can be applied cross-modally to audio encryption, thereby providing a new perspective for the design and evaluation of audio encryption schemes and facilitating integrated multimedia secure communication under a unified cryptographic and analytical framework.
Ni et al. (Fri,) studied this question.