Automatic speech recognition (ASR) technology can revolutionize screening for hearing loss, given the limited resources for clinical care and increasing older adult populations. However, existing ASR systems are designed to optimally work only for native-speaker, healthy, and relatively young adults. Through the ATHENA (Automating hearing assessments for all) project, we couple two freely available ASR systems (Kaldi-NL and Whisper) to automate four traditional speech audiometry tests: speech in quiet, speech in noise, speech in speech masker, and vocal emotion recognition assessment. We then evaluate the ASR systems performance (as the impact of decoding errors in the test results) and highlight what the customization the ASR systems should include to improve the usability by a large number of user groups: children, older adults (of varying cognitive status), non-native, and hearing-impaired individuals. As a second step, the ASR customization is explored, including the post-processing of the output transcript (e.g., removing utterances or non-test related words), limiting the ASR lexicon per test (to words related to or included in the test speech material), fine-tuning and data augmentation. Finally, the customized ASR systems are evaluated, from the aspects of the system robustness, reliability, and accessibility, for both clinical and non-clinical use.
Illan et al. (Wed,) studied this question.