Over the last decade significant advances have been made in the assessment of clinical competence of healthcare professionals.1–3 Historically, one major stimulus to these advances was the publication, over 20 years ago, of a paper by Harden and Gleeson4 describing how the long case was an inefficient method of assessment and could usefully be supplemented by the objective structured clinical examination (OSCE). Subsequently, simulated patients were shown to be a feasible method of controlling many testing situations.5 Expertise was developing in both the USA and in Europe on using new numerical techniques, such as generalizability theory and Rasch scaling, to improve the precision of measurement of human capacities.6 The First Cambridge Conference, in 1984, catalysed the synthesis of traditional (judgement oriented) and psychometric testing procedures, one portentous example being the embryonic idea behind key features testing.7 In sum, these events opened the way to a renewed interest in the psychometrics of clinical competence. This work was enthusiastically developed in the USA, the Netherlands, Canada and Australia. However, a recent survey8 and observations made by the General Medical Council during its round of visits to medical schools in the late 1990s9 highlighted continuing deficiencies in the technical aspects of examinations in the UK. Conversations with colleagues worldwide suggest that, although there are centres of excellence and much good practice, the situation in many countries is not unlike that in the UK; some medical schools, boards and postgraduate institutions have forged ahead with developments in assessment techniques while others stilllag far behind. Inconsistency or incoherence is common. Recently, even the British Medical Journal has shown bipolar tendencies, with two editorials on ‘the long case’, giving contrasting perspectives, with neither quoting or acknowledging the other.10,11 There is clearly a problem. There are reasons for this dissonance. One is philosophical. There is a sizeable cohort of medical educators who have been nurtured in paradigms at odds with the ‘technical rational’ or statistical ones that drive much of psychological measurement.12–14 Adult learning principles say a lot about negotiated goals and objectives, but very little about absolute professional standards.15 In an analogous sense, medicine's apparent general reluctance to deal with evidence-based medicine seems to be as much due to the profession's unease with statistical reasoning as it is to difficulties in dealing with the application of clinical guidelines to the patient in the chair, and the apparent disenfranchisement this represents.16 Clinicians attempting to understand the process of clinical decision making have stressed observation, experience, inference and judgement,17 and have only recently been concerned with data.Another reason may simply be ignorance on the part of those charged with devising and administering examinations. Examples of the difficulty examination boards have in reconciling conflicting paradigms can be seen in the reluctance to give up ‘negative marking’ or ‘correction for guessing’ of MCQ questions, and in bizarre and opaque procedures such as ‘close marking’. The latter arose from the difficulty of combining scores from different tests with widely variant origins or standard deviations, and from the misguided belief that attempting to make fewer fine distinctions resolved the issue rather than simply hiding, or worse still, compounding it. Another factor is the place of assessment in the educational framework – often seen as a bolt-on afterthought rather than an intrinsic part of the curriculum. Some of these problems are insoluble or very context-specific; selecting 5% ofhouse officers for advanced training in neurosurgery is a different assessment problem, requiring different strategies and solutions, from revalidating an entire profession. However alldeserve to be tackled at least by informed decision making, if not by precise algorithms. After all, assessment is probably the area of medical educationthat has the most robust evidence base. Starting in this issue, the Metric of medical education series will attempt to enable educators to better understand the basic concepts of assessment, and thus to offer an enhanced rationale for attempting measurement of clinical competence. In addition, because of the close relationship between assessment and the development of appropriate outcome measures for intervention studies, the series will also include a ‘how to’ guide on designing an educational experiment. In this latter role, the Metric series will link with another forthcoming series on research methods. Metric will comprise short, commissioned articles for the medical teacher, researcher and curriculum designer, written by clinicians and assessment specialists working together. The emphasis is on presenting the basic framework of assessment and specific psychometric or assessment techniques. The first 3 articles will cover theory of assessment, generalizability theory and study design. Thereafter the planned sequence is to cover techniques appropriate to the measurement of competence, followed by those recently developed or enhanced for looking at performance. This will be followed by articles on standards. We are asking for reader feedback on the Medical Education website (http://www.mededuc.com) during and after the Metric series: we would like yourcomments on issues raised, any controversies, and those things you always wanted to know but were afraid to ask. If there is a demand for topics that we have not previously covered, then we will try to include them. In addition, if you think you have a topic that you could usefully contribute to the series please email us on brian. jolly@med.monash.edu.auor j.a.spencer@newcastle.ac.uk.
No takes yet. Share an insight, caveat, or question.
Jolly et al. (2002) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: