Speaking proficiency is a critical factor in second language (L2) research, since it affects segmental and suprasegmental aspects of speech. However, obtaining proficiency measures often relies on standardized tests or subjective human ratings, both of which can be costly, time-consuming, or unavailable. This study investigates whether an automatic speaking proficiency assessment system can serve as a reliable substitute for human ratings in L2 research. Using multi-input, multi-output deep learning models, we examine whether the relationship between speaking proficiency and segment production, previously demonstrated with human-rated proficiency, can be replicated with machine-predicted scores. Results show that the automated system effectively mirrors native listener judgments, accurately capturing variability in L2 production. These findings suggest that automated proficiency assessments can reduce reliance on human annotation, offering a scalable and efficient tool for streamlining L2 experimental workflows.
S. Park (2025) studied this question.