Building an evidence base is analogous to laying a floor–on the one hand you could cover the terrain with large carefully interlocking research studies rather like laminated flooring, on the other hand you could painstakingly piece together a myriad of service evaluations like a Roman mosaic. For an international collaboration such as the Cochrane Collaboration, with the entire global health system to cover, the latter ‘piecemeal’ approach is clearly unsatisfactory. Indeed the challenges faced by the Collaboration, both in terms of the number of questions yet to be addressed and the aspiration to update systematic reviews every 2 years are already almost insurmountable.1 But what of library and information practice, whether specifically within health or more generally? Should we await those ‘ideal’ large studies required to resolve the profession’s longstanding questions or could we settle for more incremental evidence building? Of course library and information practice is not the only sector to debate what constitutes “evidence”. Few professions share the favoured position of medicine in having approaching half a million randomized controlled trials. Indeed, even within healthcare there is increasing recognition that there are more types of question than just ‘effectiveness’ and there is equally a place for more exotic designs such as ‘interrupted time series’ or ‘comparative studies with historical controls’.2 While unfavourable comparisons with healthcare risk provoking an inferiority complex, the field of social care provides a much more realistic benchmark. In particular, the focus on users or clients and, indeed, on client-centred evaluations is a model that sits more comfortably within our own user-focused sphere. Within the domain of evidence-based policy, there is recognition that: ‘Evaluation is important for determining the extent to which a policy has met or is meeting its objectives and that those intended to benefit have done so’.3 Methods of evaluation include both experimental and quasi-experimental, and indeed formative (process evaluation—i.e. how to improve services) and summative (outcome evaluation—i.e. whether services achieve their goals) approaches.4 Within the UK National Health Service, the distinction between evaluation and research has become increasingly critical over recent years. The principal reason for this is not some professional epiphany triggering greater awareness. Instead this distinction reflects a pragmatic response to the perceived bureaucratic red tape of requirements for research governance.5 Put simply a research proposal requires health service governance approval, whereas a service evaluation does not. Faced with this further impediment to the already pressured task of completing a work-based dissertation project within 12 months health librarians, and indeed related professionals on courses such as our own University’s Masters in Health Informatics, seek to get their projects adopted as ‘service evaluation’ rather than conduct them under the more demanding marque of ‘research’. Such pragmatism is shared by colleagues completing work-based projects for other than educational reasons. In educational terms this migration from research to service evaluation has not been a mortal blow to evidence-based information practice. Most mechanisms for conducting service evaluation are shared with research. Thus, a survey, questionnaire, focus group or other means of data collection requires the same careful deliberation, design and administration whether intended for research or for service evaluation or audit. Furthermore, perceived constraints on primary research have had useful, and presumably unintended, effects in stimulating a full variety of types of secondary data collection and analysis, ranging from systematic reviews to exploration of routine data sets. Finally, just as it might be argued that the scientific cause of laboratory medicine is little advanced by an endless succession of cohorts dissecting hapless frogs or rats, it is unlikely that the evidence base for health information, or health services research more generally, is significantly furthered by a plethora of student questionnaires targeting health service staff. So where exactly is the problem? A culture of evaluation in preference to research impacts at every stage of the investigation process. First, it determines the selection of the question for investigation; typically resulting in a more modest descriptive question rather than a more ambitious analytical one. So, for example, one might seek to answer the ‘worthy’ question of ‘What are the attitudes of staff to introduction of a clinical librarian role?’ rather than the ‘Holy Grail question’6: ‘What impact does introduction of a clinical librarian role have on the quality of patient care?’ This latter type of ‘cause and effect question is central to discussions on the funding of new outreach positions. Second, a shift away from research projects can lead to a high prevalence of ‘ahypothetical’ investigations; that is projects that lack a hypothesis. Of course, this need not necessarily be the case. It is quite possible to advance and explore a hypothesis using observational or qualitative research methods. Nevertheless, it is becoming more common to see students suggest hypotheses only at the point of exploring the data they have collected, often finding the data unsatisfactory for their newly identified purpose. The benefits of a priori specification of a hypothesis and subsequent identification of requirements for data collection are thus potentially lost to an increasing number of students. The alternative scenario is that dissertations and research projects frequently identify hypotheses for further exploration as the outcome of the entire process instead of as its starting point. If an increasing supply of questions is generated without a corresponding increase in ‘consumption’ (or resolution), this will lay down an unsatisfactory legacy for successive generations of would-be evidence-based practitioners. Third, a seemingly technical, but no less critical, point relates to statistical methods. Again a shift in emphasis results in favouring descriptive statistics (for example, basic counts and frequencies) over analytical statistics.7 At most such types of investigation can explore and establish association (that two happenings frequently co-occur) rather than causation (that one happening necessarily leads to the other).8 As a consequence, the profession may become disproportionately endowed with those who can handle only elementary aspects of data at the expense not only of their own research skills but also of their ability to interpret the research endeavours of others. Fourth, there is the purpose of research: ‘research has as its primary goal the advancement of knowledge or “Truth”. It strives to advance and extend knowledge’.9 Generalizability has thus become a significant consideration in the pursuit of research, certainly when advocated by commissioners of practically targeted health services research. Ironically many measures aimed at increasing the rigour (internal validity) of research, such as ensuring a tightly defined population within a carefully controlled environment, have an adverse effect on its applicability to the general population (external validity).10 By contrast, evaluation sidesteps such methodological debates, privileging the practical usefulness of findings to the particular service being evaluated. The above should not be interpreted as a blanket espousal of the virtues of research in contrast to those of evaluation. Elsewhere, within this journal11 and through other channels,12 I have championed the value of service evaluations. However, clearly we have much to lose if we focus only on one drawer of our evidence-based practice toolbox. What is necessary is a broad understanding of ‘fitness for purpose’, that is we do not expect evaluation to carry the load of generalizability upon its shoulders that is best borne by research. Neither should we require that a service evaluation for a clear and identifiable purpose ticks all the boxes for a rigorous research proposal. Clearly, it is important that the criteria against which each is measured are clear, explicit and distinguishable. Recent experience illustrates the importance of such a distinction. A manuscript describing an evaluation submitted to a well-regarded journal garnered reviewer criticisms relating to the absence of a hypothesis and the non-use of inferential statistics. As this evaluation did not correspond to the exacting standards of rigorously designed research, the reviewer went on to describe its design as ‘flawed’. As we pointedly remarked in our response, such a judgement ‘is like criticizing a milk float for not being a Ferrari!’ Of course, the particular contribution of an ‘evaluation’ relates to one specific area of great interest to the evidence-based practitioner. ‘Evaluation’, as Scriven reminds us, relates to the ‘valuing of something’.13 In the tripartite model that constitutes evidence-based practice, the ‘values’ of user perspectives and the ‘values’ of practitioner observations are considered essential complements to results from rigorous research. In designing or modifying services for our users, we need to evaluate what is important to them. In evaluating the success or otherwise of those same services, we need to ensure that we are not simply measuring what is most easily measured—that we are actually ‘counting what counts’.6 Finally, in designing research we need to have a clear picture of those outcomes that are considered most important by users and potential users of an intervention and, what is more, to ensure that these are included as primary outcomes. Evaluation, as Guba clearly states, is concerned with devising and testing some practical solution to one or more operating problems.14 As such it must be considered fundamental to evidence-based library and information practice. How do we tackle the acknowledged limitation that evaluation is for ‘specification’ (i.e. the investigation of a specific service), while research aims for ‘generalization’?15 While some still claim that the purpose of, and results of, evaluation should not and cannot be generalized, this fails to recognize recent advances in research synthesis. Such techniques as meta-ethnography were initially developed to examine common themes across multiple school inspection reports.16 While purists might argue that such reports were never intended to be synthesized and used to produce generalizable truths, one could say the same of randomized controlled trials which were not originally designed to be incorporated into meta-analyses. Yet, as readers know well, synthesized meta-analyses of randomized controlled trials are now plentiful and well regarded. Surely then, as long as appropriate methods for synthesis of evaluation reports are developed and utilized, it is possible to build an evidence base, and a generalizable one at that, from the fragmentary mosaic of service evaluations. Our own experience,11 and that of others,17 demonstrate the value of such a synthesis even from small sets of simultaneously evaluated outreach services. If, we only had ‘world enough and time’18 (to which we might add the necessary modern-day ingredient of money!), there is little doubt that the cause of evidence library and information practice would be considerably advanced by commissioning large numbers of well-powered randomized controlled trials. However, the ideal must not be allowed to become the enemy of the good. Provided that an evaluation report is ‘trustworthy’—however, the reader might define such—it may be used to inform decisions at a local level. In fact, rather than pursue the objective ‘truth’ implied by the term ‘generalizability’, we might do better to consider the more subjective concept of ‘applicability’—like beauty, frequently in the eye of the beholder. Furthermore, provided that a service we supply does no harm, is acceptable to users and is universally considered ‘no worse’ than competing alternatives, it does not seem useful to be overly preoccupied with demonstrating its superiority through the preferred randomized controlled trial design. The memorial conference for the distinguished librarian and scholar Leslie Morton borrowed from Isaac Newton in being entitled ‘Standing on the Shoulders of Giants’19—no doubt politically incorrect in an apparently enlightened 21st century! Faced with the prospect of a synthesized evidence base painstakingly assembled from a myriad of rigorously compiled evaluation reports, we may well have cause to appreciate the value of the converse, namely ‘Standing on the Shoulders of the Vertically Challenged’!
No takes yet. Share an insight, caveat, or question.
Andrew Booth (2009) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: