This study investigates how the intelligibility of Turkish learners of English as a foreign language (EFL) is assessed by both human raters and artificial intelligence (AI), specifically ChatGPT-4, across two different speaking tasks: a controlled read-aloud passage and a spontaneous picture-description task. Drawing on intelligibility-focused pronunciation research, the study aims to explore how task type, rater type, and pronunciation features (segmental and suprasegmental) affect intelligibility ratings. 30 intermediate-level Turkish learners of English completed both tasks, and their recordings were evaluated by three native English-speaking human raters and an AI model. Quantitative results showed that spontaneous speech received higher intelligibility scores than the read-aloud task, despite including more segmental errors. Suprasegmental features such as rhythm, stress, and phrasing played a greater role in determining intelligibility across tasks. While AI ratings closely matched human judgments in most cases, discrepancies emerged, particularly in samples where prosodic nuances were critical. Qualitative analysis further revealed that both rater types frequently flagged vowel distortions, stress misplacement, and a lack of rhythmic cohesion as common intelligibility detractors. The findings underscore the importance of integrating suprasegmental instruction in EFL pronunciation pedagogy and highlight the potential role of AI tools in supporting intelligibility assessment. Nevertheless, when it comes to assessing natural, prosody-rich speech, human judgment is still crucial. By providing empirical evidence from a Turkish EFL environment and linking L2 pronunciation research with emerging technology, this study contributes to applied linguistics.
Canan Deveci (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: