Recent advances in deep learning have enabled high-quality and high-fidelity voice conversion (VC). However, cross-gender VC remains challenging, as it requires more substantial modifications to acoustic features than same-gender VC. In this study, we analyzed the dynamic characteristics of conversion errors in mainstream DNN-based systems for cross-gender VC, focusing on two representative models: RVC and StarGANv2-VC. We aligned the converted utterances with the original target utterances in time, and computed time-series trajectories and distributions of conversion errors for three evaluation metrics: Mel-cepstral distortion (MCD), fundamental frequency (fo) error, and dynamic feature error. Additionally, we examined the correlation between these metrics and the f o difference between the input and target voices. The evaluation results showed that both systems exhibited similar tendencies across the metrics. In male-to-female conversion, the f o difference exhibited a moderately strong positive correlation with f o error, whereas in female-to-male conversion, the correlation was weak. The other metrics showed no meaningful correlation with the f o difference. Building on these findings, we plan to conduct a more detailed quantitative analysis of multiple acoustic features, which include f o and prosodic features, with a particular focus on male-to-female conversion for future model improvements.
Matsuhisa et al. (Wed,) studied this question.