Key points are not available for this paper at this time.
Student teachers can seek help from large language models when writing their academic texts during their initial teacher training, and teacher educators can use large language models to evaluate the submitted academic texts. However, existing research presents inconsistent evidence regarding whether large language models assess academic work in ways comparable to human evaluators. To our knowledge, few studies have examined evaluations made by both student teachers and teacher educators alongside those generated by large language models. This study addresses two research questions concerning how academic texts are evaluated by student teachers, teacher educators, and large language models. First, we found that the two large language models showed agreement with each other but did not consistently align with the evaluations provided by either the student teachers or the teacher educator. Second, the large language models produced substantially longer evaluation texts that closely followed the structure of the assessment criteria but struggled with evaluating the discussion sections. Although the large language models offered practical suggestions for improving academic texts, their feedback did not emphasize the same aspects highlighted by the teacher educator. Implications for practical use of generative AI-tools and needs for further research are discussed.
Gillespie et al. (Sat,) studied this question.