Model evaluation demonstrates effective meaning-level fake speech detection using multimodal gated fusion, highlighting the primary predictive value of transcript and emotion-alignment cues.
Fake-speech detection is commonly studied through acoustic artifacts, speaker-level spoofing cues, or visual inconsistencies, while less attention has been given to meaning-level manipulation, where the spoken content is altered while the speech remains natural and speaker-consistent. This study investigates the relative and complementary contributions of interpretable speech-centered features for detecting meaning-level fake speech. Specifically, it examines transcript-level linguistic and psycholinguistic-style features, text–audio emotion-alignment features, and rhythm-based audio descriptors. A feature-aware gated-fusion framework is used to analyze and combine these three feature groups, with separate branches encoding each feature type and learned branch-level weights adaptively controlling their contributions to binary classification. The framework was evaluated on FakeSpeech+, an audio-only dataset designed for meaning-level manipulation, using a strict leakage-controlled repeated-seed protocol that prevents source-pair, filepath, exact-transcript, and combined group overlap across training, validation, and test partitions. The gated-fusion model achieved 0.842 accuracy, a 0.840 F1-score, and 0.919 AUC. Analysis of the learned fusion weights indicated that transcript-level features contributed most strongly, followed by text–audio emotion-alignment features, while rhythm features received the lowest contribution. These findings provide evidence that meaning-level fake-speech detection can benefit from jointly examining linguistic content, emotional alignment, and rhythmic characteristics, while also highlighting differences in the relative contributions of these feature groups.
No takes yet. Share an insight, caveat, or question.
Alsaeedi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: