Key points are not available for this paper at this time.
Debates over teacher evaluation and teacher effectiveness have largely focused on methodological refinement, particularly the technical adequacy of value-added models (VAMs). Far less attention has been paid to the validity of the evidentiary foundations on which these systems rest. In this theoretical synthesis of validity theory related to use, interpretation, and consequences, we argue that validity must be treated not as a technical property of measures alone, but as an interpretive, political, and governance-laden construct that precedes legitimate policy use. We examine achievement tests as the primary evidentiary base for teacher effectiveness claims, highlighting persistent misalignments between test purposes, interpretations, and high-stakes uses. We then show how VAMs extend rather than resolve these foundational validity problems by amplifying weak inferential links between student test scores and judgments of teacher quality. Drawing on contemporary validity theory, we emphasize the centrality of interpretation–use arguments and consequential validity in accountability contexts. We conclude by outlining what a validity-first approach to teacher evaluation would require and why unresolved validity questions represent ethical and political failures in education policy.
Amrein-Beardsley et al. (Tue,) studied this question.