This study introduces the ECAPA-TDNN Vocal Health Encoder (ECAPA-TDNN-VHE), a deep learning model designed to learn speaker-invariant, health-centric speech representations for vocal fatigue assessment. The model is trained from scratch using a supervised contrastive learning objective on speech data from over 70 speakers recorded under diverse real-world conditions, including variations in microphones, recording devices, background noise, and speaker demographics. The resulting model produces 192-dimensional embeddings that encode vocal health characteristics while suppressing speaker identity. Vocal fatigue is quantified through a continuous, geometry-based scoring formulation, where embeddings are evaluated relative to a learned centroid of healthy vocal representations. This formulation enables interpretable fatigue scoring without reliance on discrete class boundaries, supporting both analysis and longitudinal monitoring of vocal health. To facilitate reproducibility and real-world applicability, we release auralisᵥfs, an open-source Python library that operationalizes the proposed methodology. The library provides standardized audio preprocessing (16 kHz, mono, 5–10 s segments), supports common audio formats (. wav,. mp3,. m4a), and implements robust handling of edge cases encountered in practical audio recordings. Through a simple API, users can extract embeddings and compute relative vocal fatigue scores, enabling downstream research, feature extraction, and early detection of vocal strain. All components required for reproducibility are made publicly available, including trained model weights, the healthy embedding centroid, and the fatigue axis used for scoring. The dataset is multi-speaker, multi-device, language-independent, and gender-balanced, and ethical considerations were observed throughout data collection and release. This work aims to support future research in vocal health monitoring, voice-based fatigue assessment, and computational paralinguistics, while providing a scalable foundation for health-oriented speech representation learning.
Muhammad Khubaib Ahmad (Sun,) studied this question.