This presentation demonstrates acoustic measurement tools for improving speech recordings, suggesting better quality outcomes in various recording scenarios.
Accessible recorded speech resources are expanding enormously in quantity and quality. In a companion presentation (Sakakibara et al., 2025, this meeting), we propose a protocol and quality classification for recording such resources. The goal is to make the recorded resources better in each class. In this presentation, we focus on the speech sounds we encounter in everyday life situations. This presentation introduces acoustic measurement methods and tools designed to improve the recording process. Two recent developments in signal processing make these tools suitable for a wide range of recording situations. The first, “signal safeguarding” (Kawahara and Yatabe, Acoust. Sci. Technol., 2022), enables any sound (such as lecture, talk, conversation, and music) to be utilized as a test signal for acoustic measurement. The second, “giant-FFT SRC (Fast Fourier Transform, Sampling Rate Conversion, Välimäki and Bilbao, J. AES, 2023),” enables an exact, aliasing-free implementation of sampling rate conversion and provides a solid foundation. The tools provide interactive, real-time visualization of frequency response, reverberation time, signal-to-noise ratio, and direct-to-indirect sound ratio—primary factors affecting speech recording quality. These tools are available as open-source resources. [Work supported by JSPS KAKENHI Grant No. JP23K20440.]
No takes yet. Share an insight, caveat, or question.
Kawahara et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: