Authors
In 2008, the Accreditation Council for Graduate Medical Education (ACGME) announced “The Next Step in the Outcomes-Based Accreditation System,” now commonly known as the Next Accreditation System (NAS).1 The goal of NAS was to take the “next step” in transforming graduate medical education from a process focus to one of outcomes. Central to this transformation was further defining the existing core competencies into specialty-specific subcompetencies, and developing intermediate milestones, or observable knowledge, skills, and attitudes for each subcompetency that describes performance on a spectrum from novice to expert.2 In the fall of 2013, emergency medicine (EM) participated in the first wave of milestone data reporting to the ACGME.1 Although rigor is important in medical education research and scholarship, it matters most in high-stakes learner assessment, which the milestones are. We owe it to our learners to make sure that we are measuring what we intend to measure (validity) and that our measurements are consistent over time (reliability). Quoting a recent editorial, “As the early adopter of NAS, EM has set the standard for the design, alignment, and integration of milestones into the educational framework of the future.”3 EM was one of the first specialties to develop milestones through a rigorous process that included a consensus of national experts. The result was an important source of validity evidence for milestone assessment based on content.2, 4, 5 In this issue of Academic Emergency Medicine, Beeson et al.6 present an important step in validating the data produced by EM milestones. These results add to the specialty-specific milestones assessment literature by bringing together the resources and institutional will of the ACGME and American Board of Emergency Medicine (ABEM). This large psychometric analysis of data from EM Milestones is a single-event observational study. It is impressive in its inclusion of 100% of EM residents (5,805 from 162 programs), analyzing the entire data set from the third reporting cycle from October 2014 to January 2015. The findings of this study are noteworthy for showing that when clinical competency committees rate EM residents on the 23 subcompetencies there is consistency within each postgraduate year of training and stepwise improvement for each additional year and also that there are two primary and a few additional clusters of factors that the subcompetencies group into. To understand the significance of this study's findings, a brief mention of the current understanding of validity concepts is helpful. Validity is not a property of an instrument or process. Rather, it refers to the interpretations of scores from a specific instrument in a specific context. Validity should be thought of as an argument for whether a tool measures what is intended, the results of which are always a matter of degree. In the case of milestones, validity evidence refers to using milestones as summative, by committee assessment, and do not translate to using the frameworks for other assessment data. The study “Initial Validity Analysis of the Emergency Medicine Milestones” is based on Messick's construct of validity.7 It measures reliability (0.96 within each year) with Cronbach's alpha and factor analysis, which provide internal structure validity evidence. Traditionally there has been an array of various types of validity including face, content, criterion, and predictive to name a few. Messick broke from tradition by introducing a unified concept with only one type of validity, construct validity. He also identified five sources of validity evidence: content, response process, internal structure, relations to other variables, and consequences. This concept has been built upon by others to form the contemporary unified theory of validity.8-10 Beeson et al. have demonstrated outstanding reliability for the data produced by the milestone process, which is a necessary but not sufficient source of validity evidence. As the title of their work suggests, this is an initial validity analysis. Because the importance of milestones requires a high degree of confidence in assessment, it mandates more extensive evidence than settings where a lower degree of confidence is sufficient. Consequently, there is a need for additional validity evidence from multiple sources, evaluation of potential limiting bias, and defining of the appropriate role of milestones in assessment before all stakeholders can rest assured.11 As suggested by the authors of the study, future validity evidence should focus on the accuracy of milestone assessment in measuring competency. This could be accomplished by correlating milestone results with other assessments of performance. Another potential source of evidence is “consequences.” For example, if we use milestones to make decisions about remediating residents, does the remediation prove to be necessary (i.e., are the consequences of the decision “right” or “wrong”). Factors that exert nonrandom influences on scores, such as bias or extraneous information (construct-irrelevant variance), reduce accuracy and must be evaluated more closely to determine the magnitude of their effect. This is particularly important when considering common limitations of assessment, including grade inflation, haloing, lack of faculty training in assessment, and bias based on the year of training. We applaud the authors of this manuscript for their extraordinary efforts to demonstrate internal structure validity evidence regarding the EM Milestones. We also join them in the call for additional works that provide more extensive validity evidence so that we can someday soon say, “EM residents can be assured that this evaluation process has demonstrated validity and reliability; faculty can be confident that the milestones are psychometrically sound; and stakeholders can know that the milestones are a nationally standardized, objective measure of specialty-specific competency acquisition.” We thank Larry Gruppen, PhD, for his editorial review of the manuscript.
No takes yet. Share an insight, caveat, or question.
Love et al. (2015) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: