My pet hate, nurtured over almost a decade of evidence-based practice, is when seeming authorities refer to ‘the hierarchy of evidence’. In reality, there is no single authoritative hierarchy of evidence and use of the definite article (i.e. the) implies a unitarian view of the world that is increasingly alien to the spirit of evidence-based practice. If we are to believe that evidence-based library and information practice (EBLIP) brings together user-reported, librarian-observed and research-derived evidence1 then we should expect a minimum of three hierarchies of evidence. In actuality there exist many such hierarchies – as well as a compelling argument for ditching the whole flawed premise all together. It is perhaps no surprise that the concept of the hierarchy of evidence seems to be heavily influenced by the highly- hierarchical medical profession. In this context it may be seen as primarily a rhetorical artifice to ensure that ‘my evidence trumps your evidence’!2 Hierarchies of evidence appear to have been invented for several reasons. At an individual level, and when associated with systematic reviews and guideline production, the hierarchy is designed as a means to prioritise efforts in sifting through research in order to identify the best available evidence.3 Several points can be made here. First, position on the hierarchy is not automatically determined by the inherent superiority of a particular study design. Fundamental to the final decision are assumptions about study quality. A bad randomised controlled trial (RCT) is not superior to a good cohort study – we would always prefer a recognisable presence of bias to one that is both undetermined and unquantifiable. Similarly hierarchies of evidence arbitrate effectively when we have two competing studies, of good quality but different designs, competing in a head-to-head slugfest. They cannot distinguish between two equally matched studies of similar design but conflicting results. Nor can they handle a mêlée situation with three RCTs and two cohort studies ranged against four cohort studies and one RCT, much less the tag team of eight cohort studies taking on a single RCT. Finally many such hierarchies place the systematic review in pole position. However, a systematic review is basically an observational study, albeit typically conducted by two or more people, of RCTs or other studies. Is a systematic review less biased than the individual RCTs that it summarises or is it instead a compoundly flawed product that adds its own vulnerability to bias to that present in each of its component studies? Part of the problem with hierarchies of evidence is that our judgements of the usefulness of research studies are uneasily trying to reconcile two competing considerations. On the one hand, we would like our evidence to be based on rigorous study designs with a reduced susceptibility to bias – technically known as ‘internal validity’.4 On the other hand, we would like such studies to be comparable with our own setting within which we plan to use the findings – technically known as ‘external validity’ or generalisability.5 If you are a librarian in a modest-sized hospital library would you prefer evidence in the form of a multi-centre trial of university medical libraries or as a case study in a library almost identical to your own? Yes, the answer is probably ‘it depends’ but with centuries of ‘anecdote-based practice’ yet to be overturned by the sleek scientific efficiency of the evidence-based practice movement you can hardly be criticised for inclining towards the latter. Were there a ‘hierarchy of applicability’6 it would probably attach relatively little value to systematic reviews and, even, to RCTs and would focus attention on easily replicable, albeit possibly low-quality, studies conducted at a local level. You might conclude from the above that the so-called ‘hierarchy of evidence’ is anything but that. Typically I refer to it as a ‘hierarchy of effectiveness’ although technically it could be more accurately described as a ‘hierarchy of comparative internal validity’.7 Where it works best is in emphasising the superior design aspects of randomised over non-randomised, of prospective over retrospective, and of multiple cases over a single isolated case. Ironically the need to impose a ‘controlled’ environment on a research investigation is both a strength and a limitation of the RCT methodology. On the one hand, it affords a degree of protection against sources of bias and possible confounders; on the other, it provides a short-term evaluation timeframe – ideal for measuring 1 year’s training of groups of students using two different methods (e.g. face-to-face versus e-learning) but less optimal for following the effect of that training longitudinally over 5 years of post-graduation. Jonathan Eldredge was one of the first authors to discuss the hierarchy of evidence in the context of EBLIP. He distinguished the then ‘EBL’ paradigm from traditional LIS research by its focus on ‘higher levels of evidence than those traditionally used in LIS research: descriptive surveys, case studies, and other qualitative methods’.8 Eldredge first proposed a hierarchy of evidence for EBL that was virtually identical to the levels of evidence used in evidence-based medicine.9 Recognising the practical problems that such an approach poses to the EBLIP practitioner Crumley and Koufogiannakis10 identified the traditional hierarchy of evidence as a potential barrier to librarians finding research to answer their day-to-day questions. How is a health librarian to view the evidence base for their practice if the prevailing view is that ‘anything less than a randomised controlled trial is not sufficient evidence to answer their questions or change their practice’?11 Examining the state of the evidence base in 2001 they identified that only 12 of the 807 research articles published in 91 library and information studies journals could be described as systematic reviews, meta-analyses, RCTs, or controlled trials – recognised as ‘higher levels of evidence’.12 At the 4th International EBLIP Conference in North Carolina, Rowena Cullen,13 one of the keynote speakers, addressed hierarchies in the context of evidence for e-government. She highlighted how evaluation of the use of online technology in government has rapidly fallen behind the prevailing pace of the technology itself. Complementing Crumley and Koufogiannakis by focusing on the production, as opposed to the consumption, end of the evidence, Cullen suggested that the traditional hierarchy of evidence was unlikely to be relevant in this context of evaluation – or at least the methodologies at the top of the evidence hierarchy held little promise in addressing the evaluation gaps. Instead she stated that methods from the toolkits of policy or program evaluation or those associated with descriptive and formative evaluation may be more appropriate. This theme of privileging evaluation over evidence has recurred several times in the EBLIP literature over recent years.14 We might conclude that if the perpetuation of the evidence hierarchy within EBLIP is equally problematic for producers and consumers of evidence alike then what is it good for? To which it is tempting to add – absolutely nothing! What alternatives exist to the simplistic concept of an evidence hierarchy? Let us consider briefly in turn just two – signal-to-noise ratios and evidence typologies. The idea of a signal-to-noise ratio as an alternative to hierarchies of evidence was first advanced by Edwards et al.15 They pointed out that there is a need to make best use of the available evidence in answering questions that may not be addressed by clinical trials. They therefore proposed a classification of research which does not reject studies on the basis of design alone, but recognises the importance of assessing the message or ‘signal’ within each piece of research. Although fundamentally flawed research will necessarily be rejected, papers with less rigorous designs can still be used, provided that we temper the importance that we attach to their signal with the amount of ‘noise’ accompanying that signal. Methodological deficiencies increase the amount of noise while a match with the context or population in which you plan to use the evidence boosts the signal. In this way effectiveness, feasibility and appropriateness can be factored into our judgements of an article.16 In truth the concept of ‘signal to noise’ has not received the attention that we might have expected. While intuitively it feels like a useful approach it perhaps seems difficult to operationalise. Nevertheless in abstract sifting sessions with librarians attending EBLIP workshops we have been able to consider evidence across both dimensions, signal and noise, simultaneously. This combination of the ‘signal’ from available research papers and the level of ‘noise’ (the inverse of methodological quality) has been articulated as the ‘weight of evidence’ in a particular topic area.17 It has been suggested that this approach, when applied to the ‘qualitative overview’ stage of an evidence review, allows due consideration of individual items of evidence not otherwise considered important within a more quantitative analysis. For example, within the context of EBLIP, a single case study of considerations around training students to staff a reference desk – a low quality oasis in an otherwise barren evidence desert – may contribute to our holistic interpretation of the more plentiful and more rigorous literature surrounding models of reference desk staffing. EBLIP is by no means the only field to believe that the ‘hierarchy of evidence’ is difficult to apply. Clearly, however, we cannot simply ‘shelve’ the hierarchy without having some other framework to replace it. Petticrew and Roberts18 suggest a typology based around a matrix which matches research questions to specific types of research. Hansen and Rieper19 argue that a hierarchy of evidence stresses the importance of internal validity while a typology of evidence stresses the need to consider which type of studies are most appropriate for answering different types of questions (Table 1). This emphasis on typologies rather than on hierarchies of evidence was championed within EBLIP by Eldredge when he proposed his levels-of-evidence table8 to replace the hierarchical list previously suggested.9 In using this table Eldredge distinguished three categories of research questions:8 Prediction questions classically addressed by the cohort study design. Intervention questions comparing two or more alternatives to determine which is better. Exploration questions which imply a ‘why’ enquiry and which draw upon focus groups, in-depth interviewing, Delphi techniques, observation and historical analyses. The addition of qualitative research designs, specifically in answering exploration questions, resonates with thinking within wider evidence-based practice.21 Librarians must be prepared to locate, critically appraise and use other kinds of research, including qualitative research, to inform decision-making –‘finding usable research for practical situations’.22 Eldredge proposes the systematic review as the highest level of evidence for all three categories of questions8– the strength of the ‘signal’ is boosted by findings derived from multiple ‘channels’ or studies. Thus the qualitative evidence synthesis (a.k.a. qualitative systematic review) may prove a useful tool in exploration of longstanding questions around user preferences, attitudes and behaviours. We should not be surprised that the embryonic evidence hierarchy appears to have outlived its usefulness. It was very much a product of its time and reflects an earlier, less complex view of the process of evidence-based practice. While evidence typologies, in the form of matrices, are very much the flavour of the present, it is likely that, in time, they too will be revealed as an oversimplification. Perhaps, mirroring developments elsewhere, the next manifestation of a tool for systematic assessment of evidence will be the ‘evidence network’. Might we conceive that all the necessary dimensions of evidence, e.g. appropriateness, effectiveness and feasibility, might be mapped on a spider (radar) chart23 and then we examine the extent to which this multiform shape corresponds to the shape of our originating question? In this way the evidence that most closely matches our purpose will be privileged over other forms and we can make explicit trade-offs between user views, effectiveness and practitioner observations as well as other considerations. Of course use of a composite approach to rating evidence will make testing demands on our need to communicate this approach effectively. We would do well to remember the so-called Law of Communication: ‘The inevitable result of improved and enlarged communications between different levels in a hierarchy is a vastly increased area of misunderstanding’!24
No takes yet. Share an insight, caveat, or question.
Andrew Booth (2010) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: