Abstract : Artificial intelligence (AI) ethics has emerged as a global focus in K-12 education, with a shared goal of nurturing informed and responsible citizens in AI-infused societies. However, existing instruments of students’ ethical development often rely on self-reported tools, with few objective tests. This study addresses this gap by developing and validating a scenario-based multiple-choice test for introductory AI ethics education among early adolescents. Following a theory-driven approach, this study operationalizes AI ethical competence (AIEC) as a multidimensional construct defined by students’ abilities to recognize ethical issues, reason through consequences, and take appropriate actions. The instrument underwent rigorous development, including cognitive interviews ( N = 24), expert review, pilot testing ( N = 289), and field testing ( N = 5777; mean age = 12.37). Psychometric evaluation was conducted using reliability analyses and Classical Test Theory (CTT) and Item Response Theory (IRT) analyses. A final 17-item version demonstrated good item-level performance, with moderate difficulty ( b = −1.440 to 0.774) and discrimination ( a = 0.629 to 2.052). The Person-Item Map showed strong alignment between item difficulties and student abilities for the central 70–80% of the sample. DIF by gender revealed a content-based pattern: all items with large DIF involved solution-oriented content and utilitarian reasoning, suggesting gender differences in response patterns for this type of task. This study provides a novel, psychometrically validated test of AIEC in K-12 education. The instrument supports diagnostic and formative assessment in school settings, while laying a foundation of refined tools for summative assessment. Additionally, the scenario-based format highlights the sociotechnical nature of students’ ethical development and calls for culturally relevant approaches to AI ethics education. • Developed a scenario-based multiple-choice test to assess students’ AIEC • Used item response theory (IRT) to evaluate item and test-level performance • Final 17-item test showed acceptable reliability and moderate item discrimination • No DIF by school level; gender-based DIF found in solution-oriented items • First validated tool for K-12-focused objective test in AI ethics education
Dai et al. (2026) studied this question.