Artificial intelligence (AI) has become increasingly embedded in technology-enhanced learning environments, where peer assessment now serves both instructional and analytic purposes. Beyond allocating feedback and grades, it also produces data that is later interpreted through learning analytics systems. In practice, visible indicators such as students’ fairness perceptions and the degree of agreement among peer raters are often treated as signs that the assessment process is functioning effectively. However, these indicators do not necessarily correspond to grading validity. Students may regard a peer assessment process as fair even when peer-generated ratings remain weakly aligned with expert judgement. This study, therefore, examines whether the socio-technical configurations associated with high perceived fairness in a digitally mediated peer assessment environment also correspond to criterion-referenced grading validity. Data were collected from 215 undergraduate students enrolled in an Artificial Intelligence Foundations course over two consecutive semesters at a university in Taiwan, with instructor ratings serving as an external expert reference within the course context, rather than as a universal ground truth. Because anonymity conditions and semester were fully confounded in the study design, differences linked to anonymity should not be interpreted as isolated causal effects. A two-stage fuzzy-set Qualitative Comparative Analysis (fsQCA) was used. In the first stage, three equifinal configurations associated with high perceived fairness were identified. In the second stage, these configurations were examined against four grading objectivity outcomes: peer–instructor alignment, peer convergence, familiarity bias, and leniency bias. The findings show that fairness perception and grading validity are only partially aligned. Configurations anchored in explicit criterion transparency consistently supported both experiential legitimacy and evaluative accuracy. By contrast, one configuration was associated with high peer convergence while remaining weakly aligned with instructor standards, a pattern described here as false objectivity; this context-dependent configurational finding warrants further investigation across other settings. The study contributes to research on digitally enhanced assessment and learning analytics by showing that fairness perception, peer convergence, and grading validity should be treated as analytically distinct dimensions of assessment quality.
Huang et al. (Sat,) studied this question.