• The proposed SWF tokenization method, dynamically capturing emotional subtleties, noise robustness, and contextual information to robustly preserve speaker identity. • A transformative deep-self-supervised spectrogram transformer back-end, outperforming conventional approaches by effectively addressing their inherent limitations in preserving local features without the need for extensive labeled training data. • Complete elimination of dependency on additional speech enhancement, enabling seamless, efficient, and robust end-to-end learning tailored for real-world deployments. Identifying speakers in noisy and emotional conditions remains a significant challenge due to the distortion of spectral cues. This study proposes the Speech Without Filter (SWF) framework, a novel self-supervised learning paradigm that operates directly on raw spectrograms. Theoretically, this research introduces a progressive tokenization mechanism that acts as a structural inductive bias, mimicking the contracting path of a U-Net to preserve local spectro-temporal continuity. Unlike standard fixed-patch Transformers that often smooth over speaker-specific micro-textures, the SWF architecture integrates denoising and feature extraction into a single stage, challenging the traditional decoupled paradigm of speech enhancement and recognition. Using a sample of 1.58 million pre-training instances, the model was evaluated across English (RAVDESS), Arabic (ESD), and stressful (SUSAS) datasets. Results demonstrate significant improvements, with the SWF model achieving 91.01% accuracy in clean conditions and maintaining 88.5% in high-noise cocktail party environments, outperforming state-of-the-art models like WavLM and HuBERT. These findings suggest that architectural innovation in tokenization is as critical as pre-training scale for robust speech processing.
Hamsa et al. (Fri,) studied this question.