The Head-Related Impulse Response (HRIR), which characterizes the sound waves reaching the eardrums and contains essential information for accurate sound localization, plays a vital role in spatial audio synthesis. Because HRIRs are strongly influenced by individual anthropometric features—such as the size and shape of the head and ears—obtaining personalized HRIRs is highly desirable. Typically, HRIRs are measured by recording responses from multiple directions using a loudspeaker array. However, as it is impractical to perform such measurements for every individual, research has focused on estimating personalized HRIRs. Nevertheless, improving the estimation accuracy remains a significant challenge. In recent years, various neural network-based approaches have been proposed to estimate individualized HRIRs. In this study, we aim to enhance estimation accuracy by employing Deep Operator Networks (DeepONets). DeepONets are capable of flexibly learning mappings conditioned on inputs by separately modeling condition and coordinate spaces. In our experimental setup, the direction of the HRIR is provided as coordinate input, and the individual's anthropometric features as conditional input; these are processed through separate subnetworks to directly estimate personalized HRIRs.
Ikeda et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: