In the previous issue, Stuardi et al.1 described the potential advantages and problems with using primary care databases to identify possible trial participants. This is yet another example of the richness of material stored in primary care records, and how it can be tapped into for use in research. We address the use of electronic databases for research here. The electronic recording of clinical patient data has come a long way since its inception during the 1980s. What started out as one UK-based database of patient-related diagnoses, prescriptions and demographic information has rapidly expanded into multiple national systems. In the UK alone, there are three predominant systems used both for research and for clinical records: the General Practice Research Database (GPRD), the Health Improvement Network (THIN) and QResearch. For all three systems, the information relating to symptoms, diseases, consultations and other clinical events are categorized using the Read code system. This structures its clinical data on the ICD-9-CM (International Classification of Diseases, Ninth Edition, Clinical Modification); Read codes are however rarely used outside the UK. Mainland European multi-practice databases also exist, as does a US-based Veterans’ healthcare system (US Department of Veterans’ Affairs—VA). They contain coded material using the more popular International Classification of Primary Care (ICPC) coding system. Less research has originated from these databases. The focus on using electronic patient records (EPRs) increased from the 1990s, accelerating recently as a result of both UK government initiatives on providing incentives for GPs to develop electronic clinical registers (i.e. the Quality and Outcomes Framework) and through the proposed implementation of a single linked National Health Service (NHS) computer system of electronic patient health records (National Programme for Information Technology, NPfIT). The UK established an early reputation for producing quality EPR-based research through the publication of numerous studies assessing the validity and quality of both the databases themselves and their resulting epidemiological research. These articles have confirmed the high quality of data recorded in primary care records, particularly relating to chronic disease diagnosis and prescription data.2–4 There remain areas for improvement in coding practice5 making use of data recorded outside the traditional fields (the ‘free-text’ data)2 and of secondary care and laboratory data.6 These last two have been mitigated to some extent by recent linkage to cancer registry data and to hospital episode statistics. Furthermore, most UK practices now receive laboratory results electronically, allowing studies to be based solely on such results.7 The advantages of basing research on a national system of networked data are well documented. Firstly, the use of medical data collected by individual practices across the UK provides not only representation of the local population but the geographical spread of the databases also allows for broad generalization to the UK population. Secondly, many patients have complete data available from the 1980s, allowing for analysis of health over a patient’s lifetime and providing a good source of longitudinal data for the researcher. Thirdly, the size of the databases allows research on low incidence diseases logistically impossible to study in any other way. The GPRD database alone currently contains research worthy data on ∼5 million patients from just under 600 primary care practices from the UK (GPRD website: http://www.gprd.com/home/), with a similar number of practices providing data on ∼13 million patients in QResearch’s EMIS database (QResearch website: http://www.qresearch.org/SitePages/What%20Is%20QResearch.aspx). The automatic recording of all this data provides quick access for researchers and reduces the time and expense of recruiting target populations and searching for controls. This also means that research targets specific vulnerable populations that researchers ordinarily may have trouble recruiting, such as psychiatric patients. A recent study (albeit from a community-based mental health database, as opposed to a primary care one) assessed the impact of the 2009 swine flu pandemic on children and adults with mental health problems in south London,8 showing the disproportionate anxiety arising from the pandemic on this group of vulnerable patients. The accuracy of research findings depends on the quality of information entered into the patient record. One inescapable problem is that data are only available from those patients who chose to consult their GP. Additionally, in some fields, the recording of data is either insufficient, absent or unavailable for use. Ethnicity is rarely recorded. Furthermore, in contrast to a traditional cohort study, whereby the population can have regular measurements of the variable of interest recorded (such as weight), in EPRs data entry is entered sporadically. This is partially ameliorated when there are incentives for systematic recording. Much data from other health care sectors are sent to GP practices in paper form, requiring each practice to input the data on to the patient record. Given how labour intensive it would be to import all this data, it is inevitable that some information is lost as the correspondence itself is not available to researchers. In theory, UK centralization of medical records could have eliminated this problem, though at present, this possibility is likely to be thwarted. For the information that is available in research EPRs, there is a hierarchy of accuracy, with the best being prescription data, with decreasing quality in diseases and lifestyle/socio-economic data.2,4,5,8 It is hardly surprising that the recording of prescription data is so good, given that prescriptions are computer generated and that pharmacoepidemiology made up a large proportion of early EPR-based research and was the original rationale for establishing the GPRD. The poorer recording of lifestyle parameters such as exercise or diet may reflect either a perceived lower clinical importance or difficulties in categorization. The quality of information in EPR databases has been kept high by imposing data entry standards on participating practices. What is problematic, however, is how each GP codes their data (which can be idiosyncratic); indeed, GPs can derive their own Read codes for their individual practice. In one study, >25 different codes related to diabetes were in use across 17 GP practices in one area of London.5 This poses real problems for researchers and national registries alike as the use of more obscure codes (or codes solely relating to treatments/prescriptions for the disease in question) may be overlooked when selecting cases for research inclusion or in compiling official statistics.3,5 Fortunately, most obscure codes are used infrequently. There is considerable inter-practice variation in coding: the generic diabetes Read code, C10 was used across GP practices from the above study to classify diabetes in between 14% and 98% of diabetic patients.5 Similarly, a study of cancer diagnoses in EPRs, uncovered several false positives and false negatives likely to have been caused by transcription errors.10 In the area of cancer, this situation may have improved, probably reflecting financial incentives to improve UK cancer recording.11 A systematic review showed good concordance between the GPRD and the Morbidity Statistics from General Practice 1991–92 (MSGP4) on most diseases and chronic conditions but lower GPRD prevalence rates in diabetes, musculoskeletal disorders and several other conditions.2 Suggested improvements to the quality of the data include extrapolating information from the free text area of data recording which may contain additional information related to a disease diagnosis;3 secondly, there is also a definite need for defined coding rules to aid consistency of data recording.12 Aside from the issues relating to data recording, researchers must weigh up the costs associated with accessing patient records (as determined by each EPR database) with the convenience of large datasets readily available for use. The charges involved in obtaining a database are almost always less than would be required to obtain such data conventionally. However, researchers require powerful computer systems to undertake studies on EPR data; plus considerable skill is needed for data manipulation as well as clinical knowledge for interpretation. It is always easier to spot disadvantages than advantages. Electronic database research is mushrooming and not just because the data are readily available. This is because of the size, generalizability and richness of the data; however, researchers need to remember the limitations and where possible sidestep them. Many high-quality studies have been performed using EPR databases, with the first phase being pharmacoepidemiology. Phase two—of disease epidemiology—has begun and is potentially even more exciting.
No takes yet. Share an insight, caveat, or question.
Shephard et al. (2011) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: