Key points are not available for this paper at this time.
This study introduces an artificial intelligence–human-in-the-loop (AI-HITL) process to create an instrument for evaluating the quality of learning outcomes (LOs) in a veterinary curriculum. Although clear LOs are essential for competence-based education, their quality often varies, and traditional manual review is inconsistent and time-consuming. We used a large language model (ChatGPT-4o) to generate the initial 12-item Learning Outcomes Evaluation Instrument (LOEI) intended to measure clarity, measurability, and alignment. This instrument was then used by human experts to rate 415 LOs. Psychometric analyses supported a strong, unidimensional structure for the reduced to four-item LOEI, with high inter-item correlations ( r = 0.53 to 0.97), excellent model fit (CFI = 0.998, TLI = 0.995), and strong reliability (ordinal α = 0.92, ω = 0.93). These findings indicate that the AI-HITL–developed instrument demonstrates internal coherence and reliability, though evidence of item redundancy motivates further instrument refinement. We highlight the promise and limitations of using AI for instrument development and confirm that human expertise remains essential for psychometric validity and contextual relevance to improve curriculum quality and alignment with competence-based veterinary education.
Schmidt et al. (Thu,) studied this question.