PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 10, 2026Language Testing3 citations

Assessing Interactional Competence Through Generative AI: Comparing Large Language Models as AI Interlocutors in the Paired Oral Discussion Test

View Full Paper
INInyoung Na

Key Points

  • This study aims to assess the effectiveness of large language models in eliciting interactional competence during oral discussions.
  • Twelve international students completed paired discussion tasks with two large language models in counterbalanced order.
  • System performance was evaluated using metrics like breakdown activation consistency and stance maintenance.
  • Interactional discourse analysis identified features of interactional competence across three dimensions.
  • Claude outperformed GPT-4o in eliciting interactional competence features and activating breakdown strategies.
  • Test takers perceived Claude as more authentic and natural compared to GPT-4o, which was seen as more artificial.
  • Different LLMs created distinct interactional conditions affecting elicitation of interactional competence and participant perceptions.

Abstract

Interactional competence (IC) is essential for oral communication assessment, yet human partner variability can introduce construct-irrelevant variance in paired speaking tests. As an alternative to a test with a human interlocutor, this study describes the development of a large language model (LLM)-driven Spoken Dialogue System and compares GPT-4o to Claude 3.5 Sonnet to inform model selection for IC assessment. Twelve international students completed paired discussion tasks with both LLMs in counterbalanced order. System performance was evaluated through breakdown activation consistency, stance maintenance, and persona adherence. Test-taker performances were analyzed using interactional discourse analysis to identify IC features across three dimensions: topic management, interactional management, and interactive listening. Semi-structured interviews explored test takers’ perceptions of the AI partners. Results showed Claude outperformed GPT-4o in eliciting IC features, successfully activating communication breakdown strategies and maintaining oppositional stance, thereby creating more opportunities for test takers to demonstrate key IC abilities. Test takers perceived Claude as more authentic and natural, while GPT was perceived as more artificial. These findings demonstrate that different LLMs create distinct interactional conditions affecting both IC elicitation and test-taker perceptions. The findings highlight the need for construct-driven evaluation criteria when selecting LLMs for language-assessment contexts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Inyoung Na (2026) studied this question.

synapsesocial.com/papers/6a508df96eeac72a437a1196https://doi.org/10.1177/02655322261458364
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Can Large Language Models Simulate Spoken Human Conversations?2025 · 5 citations
  2. 2Authenticity in language testing: some outstanding questions2000 · 28 citations
  3. 3Designing interactive, automated dialogues for L2 pragmatics learning2017 · 13 citations
  4. 4Using thematic analysis in psychology2006 · 192,673 citations
  5. 5Assessing paired orals: Raters' orientation to interaction2009 · 199 citations