Novel framework detects and describes vocal modifications in manipulated audio, enhancing media integrity.
Singing voice manipulations have become increasingly common in modern music production. While such techniques can serve as creative tools to enhance artists’ expressive possibilities, they can also raise concerns about content authenticity and media integrity. To counter potential misuse of vocal manipulation tools, recent research has developed detection systems. However, these are typically limited to binary classification, indicating only whether a vocal track has been altered, without providing any further interpretable information. In this work, we address this limitation and propose a novel framework that combines forensic audio analysis with natural language generation to both detect and describe modifications in singing voice signals. Building on recent advances in audio-language models, we construct a dataset of manipulated and synthetic vocals annotated with detailed textual annotations, which we use to train and evaluate our framework. Our approach identifies and characterises a wide range of vocal transformations, including pitch correction, pitch shifting, time stretching, and singing voice deepfake generation. Experimental results show that the proposed method not only surpasses existing baselines in classification accuracy but also provides substantially greater interpretability, as it provides explanations of the outputs in natural language, making them understandable to non-experts. This makes the system particularly relevant for music production, media forensics, and copyright verification, offering a transparent and descriptive account of vocal alterations.
No takes yet. Share an insight, caveat, or question.
Moghaddam et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: