Key points are not available for this paper at this time.
Artifact-based assessment assumes that the quality of a student's work reflects the student's underlying expertise. Large language models (LLMs) break this assumption by decoupling output quality from user knowledge, which renders artifact-only evaluation diagnostically unreliable. This shift also exposes a long-standing imbalance in engineering education, in which the profession has relied on both production and review, yet curricula and grading have emphasized production far more than the skills that make review effective. In engineering, review is the disciplined detection and correction of errors that can propagate into unsafe designs, incorrect calculations, or noncompliant decisions. Historically, teaching review at scale has been constrained by the scarcity of high-quality, realistic, and flawed artifacts. LLMs invert this constraint by generating plausible-but-wrong engineering work on demand, at controllable levels of subtlety and difficulty. Consequently, we argue that assessment validity can be strengthened by treating review performance in terms of error detection, diagnosis, and justification as a primary learning outcome rather than treating artifact production as the primary proxy for competence. We position review competence as a domain-specific form of critical thinking and evaluative judgment, anchored in the constraint structures that distinguish engineering verification from generic critique. Building on this perspective, we present a taxonomy of LLM-generated engineering errors, categorized by error type, severity, and detectability. We also provide templates for prompts and an accompanying LLM interface for generating on-demand, review-centered instructional activities.
M.Z. Naser (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: