The evaluation of complex professional practice raises persistent challenges concerning construct representation and the defensibility of judgements. Teacher preparation programmes (TPPs) provide a particularly consequential context in which to examine these challenges as evaluating teaching practices remains a contested and uneven process. This systematic review examined 11 evaluation instruments used in TPPs, analysing development, implementation, and trustworthiness. The review explored how teaching practices are judged and ways systemic complexity shaped the use and effectiveness of these tools. The analysis identified recurring concerns around reliability, rater preparation, construct representation, and the alignment of instruments with professional standards. Viewed through Social Judgement Theory and Complexity Theory, the findings demonstrate how instrument design structures the evidence available for judgement, while contextual and human factors continue to shape how that evidence is interpreted. The review contributes to educational evaluation scholarship by showing why reliability alone is insufficient for defensible high-stakes evaluation and highlighting the importance of construct representation, rater cognition, contextual validity, and consequences of use. Adaptive and modular approaches are proposed to balance comparability and accountability with the contextual and developmental nature of professional competence.
No takes yet. Share an insight, caveat, or question.
Anderson et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: