Notwithstanding the growing role of speech emotion recognition (SER) in human-computer interaction, deep-learning models often display a gender imbalance, recognizing female speech less accurately.We explore a gender-aware conformer-transformer framework to solve this, which first learns gender-related patterns under noisy conditions instead of undergoing the standard end-to-end training.These patterns are then fed into a transformer emotion decoder via cross attention.A conformer encoder was pretrained on gender classification under heavy noise with a robustness of 97.73%.This framework achieved a weighted accuracy of 88.64% and narrowed the gender gap (91.1% recognition accuracy for female speech versus 85.2% for males) on the Ryerson audio-visual database of emotional speech and song dataset.The emotion recognition dropped, as expected, on the crowd-sourced emotional multimodal actors dataset without adaptation, owing to domain differences.However, the gender classification was reasonably accurate (F1 = 0.386).We conclude that using explicit gender information improves fairness under matched conditions, although cross-dataset generalization remains challenging.
Haiyun Ma (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: