Key points are not available for this paper at this time.
The rapid development of artificial intelligence technology has significantly improved the quality of computer-synthesized voices in modern text-to-speech (TTS) engines. Various appealing attributes can be added to these synthesized voices to support their widespread use in instructional videos. However, whether such synthesized voices can replace high-quality human-recorded voices remains uncertain. We conducted an eye-tracking experiment to examine the learning outcomes of instructional videos. We compared differences in learning performance, attentional engagement, and persona perceptions between a human-recorded voice and two computer-synthesized voices (formal and cute) generated by a modern TTS engine. Thirty university students participated in this study, with their eye movements recorded and analyzed as they watched instructional videos featuring different forms of narration. Overall, no statistically significant differences were found in persona perceptions between participants who learned from the human-recorded voice and those who learned from the two synthesized voices. However, the human-recorded voice significantly improved learning performance and attentional engagement. Our results indicate that while the quality of software-generated voices has reached a relatively high level of perception, it does not positively influence learning performance and attention. Therefore, we recommend that instructional video designers prioritize human-recorded voices over software-synthesized voices.
Jing et al. (Sat,) studied this question.