The use of large language models (LLMs) in language assessment, particularly in spoken dialogue systems (SDSs) for assessing speaking, remains at an early stage. This study explored the use of an LLM-driven SDS to assess second language speaking ability. Using a within-participant design, we compared the paired discussion performance of 30 participants interacting with a self-built LLM-driven SDS (E-Talk) versus a human interlocutor, focusing on the interlocutor effect on test scores and fine-grained linguistic and interactional features. Results did not yield substantive differences in oral performance across the two conditions, although the interactions with the LLM-driven SDS displayed slightly lower scores for pronunciation and language use, slower speech rate, and reduced lexical diversity alongside increased syntactic complexity, and greater initiative in introducing new ideas, prompting responses, and guiding discussions toward negotiation, coupled with weaker interactive listening. These findings suggest that LLM-driven SDSs can serve as a potential, usable interlocutor in dialogic speaking assessments, eliciting key aspects of speaking ability. That said, further refinement is needed to better capture interactive listening and collaborative meaning construction, highlighting the importance of interpreting SDS-based performance relative to its specific interactional affordances.
Min et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: