There are many primary care studies in which the research question aims to discern the frequency of health care service utilization, or the frequency of visits to a clinical provider. Some examples include evaluating comparative use of services (1), building typologies based on utilization patterns (2) and health economic evaluations (3). Often, data on the use of health services are collected retrospectively from large administrative databases, patient medical records, self-reported health care service utilization or a combination of these strategies (4). Utilization data sourced from administrative databases and patient medical records is often considered the gold standard given that data is collected at the time of encounter with the health care system. However, in some situations, use of administrative data can be time consuming, costly, and labour intensive(4) and is not without concerns regarding quality and comprehensiveness (5). For example, use of services in the private sector may not be reflected in primarily ‘public’ data sources,(6) and services rendered by health care professionals under capitation payments may not be reflected as individual visits. When such challenges arise, self-reported data can offer a feasible and comprehensive alternative. In self-reported health care service utilization, the service user (e.g. study sample participants) directly reports on their service use. The most common method in this context is a study-specific survey that is either self-administered or administered by an interviewer (7). An example of self-administration is the Canadian Community Health Survey used by Statistics Canada to gather data on health care service utilization across the country (8). This survey is mailed to households on an annual basis and depends on a large sample of respondents who are geographically dispersed, thus necessitating a low cost, low resource option. Interviewer–administration may be more suitable in difficult to reach, complex or marginalized populations. An example is a study where accessibility and quality of health care services was evaluated in female prisoners at a detention centre through face-to-face interviews conducted by specially trained research assistants (9). Regardless of the mode of data collection, reliance on participant recall of past utilization events results in data that are prone to recall error—inaccurate or incomplete recollection—which can lead to recall bias if there is a systematic under- or over-reporting of events (10, 11). In the case of random recall error, all participants are equally affected hence precluding any important bias. However, when recall error is not random and only certain participants are affected based on some characteristic, this can lead to an important bias that can threaten the reliability of results. In this article, we present sources of recall bias in self-reported health care service utilization data, as well as measures that can be instituted to reduce such bias. We conclude with a specific example in the context of a study conducted in primary care on lung cancer diagnostic pathways. Systematic errors in self-report stem from a number of different but related causes. When asking about utilization events that occurred in the past, participants may fail to remember the event entirely, otherwise known as memory decay (12). The extent to which this error manifests is, in part, related to the length of the recall period. Although a longer period of recall (i.e. utilization over a 12-month period versus a 6-month period) can yield a greater quantity of information, the accuracy of the information decreases as the recall period increases (11). This inverse relationship means that quantity and quality must always be balanced, especially when the type of utilization activities differs among participants in the study sample. For instance, salient and less frequent utilization events, such as hospitalizations and specialist visits, have been shown to be more accurately self-reported (13, 14), compared to more typical and frequent events, such as general practitioner visits (15, 16). Therefore, if some participants in the study sample had more visits to a general practitioner and other participants had more specialist visits, event recall may not be as accurate in the former as compared to the latter. Validation studies of self-reported health care utilization have shown under-reporting to be more common than over-reporting, especially with recall periods of 12 months or more (15). For type of utilization activities, under-reporting is especially evident in the context of primary care visits (10, 12) and stigmatized visits like those related to mental health (4). Although the focus of this article is not on recall bias by design type, the extent of self-reported data accuracy, including the possibility of under- or over-reporting may also be affected by study design. For example, in case–control designs participant recall of exposure may be affected by whether or not the outcome of interest is present. Memory decay can be further exacerbated in elderly populations where cognitive ability, such as memory impairment, can lead to inaccurate reporting of events (12, 17). It should be noted, however, that the literature is somewhat conflicted on this as several studies have found no relationship between demographics and self-report accuracy (18, 19). For studies that have shown a relationship, under-reporting has been more commonly found among the elderly populations (20, 21). A recall bias related to the frequency of utilization events is rounding where events with large values are rounded up by the participant. The extent of rounding tends to increase as the number of events increase—visits up to five may be reported as raw values but then rounded by five’s up to about 20 visits, and then rounded by ten’s after that (22). This may further be the case in temporal ordering of utilization events where participants are asked to place events in order of when they occurred—1 year ago, 5 years ago, 10 years ago, etc. Frequency of events can also be overestimated in a recall bias known as forward telescoping where events that occurred before the period of interest are remembered to have occurred within the period of interest (23, 24). Conversely, events occurring after the period of interest can be reverse telescoped within the period of interest as well (12). Finally, confusion that results from ambiguous or poor-quality questions, and poor communication between the interviewer and interviewee can promote recall bias (11). These factors can be compounded by participant stress, motivation, and interview dynamics, all of which can have profound impacts on data accuracy and be quite variable within a study sample (25, 26). Recall bias cannot be eliminated and should therefore always be acknowledged in the limitations of a study that involves self-reported health care service utilization. To maintain the highest level of data accuracy, however, there are several measures that can be incorporated to minimize its impact. Bhandari and Wagner (12) present a conceptual model highlighting modifiable and fixed attributes that can affect the accuracy of self-reported data. Modifiable attributes include questionnaire/interview design, mode of data collection (e.g. phone, mail, face-to-face, online) and memory aids. In the same model, recall timeframe (i.e. recall period), utilization type and utilization frequency are also presented as modifiable; however, it can be argued that these attributes are simply a reflection of the health care utilization activities of interest and the optimal recall window (i.e. the research objective), and thus cannot be modified. Fixed attributes include cognitive and psychosocial differences within the sample that can lead to variability in the interpretation of questions presented in a survey or asked by an interviewer. The modifiable attributes offer opportunities to improve accuracy and validity of utilization data. With respect to the questionnaire/interview design, valid and reliable data can be obtained by formulating questions that are clear and precise to reduce variation in comprehension (27). In addition, backwards recall can facilitate memory recall by following an ordered sequence of events; start with the present and think backwards to a point in time (28). The more recent events are easier to recall and help with the recall of previous, less recent, events (25). The opposite of this would be forward recall where a causal sequence of events is followed; go back to a point in time and think forward to the present. Memory aids, such as personal diaries, and probes, such as follow-up questions on a specific utilization event, can also enhance recall and reduce the risk of under-reporting (29). In terms of mode of collection, in-person interviews have been speculated to lead to more accurate recall data (10, 30). As the length of the recall period can affect data accuracy, this should be considered during the development of study methods. Primary importance should be given to meeting research objectives but with a careful balance regarding the optimal recall window and data integrity. Long recall periods of 12 months or more could be justified in a study focused on inpatient visits (i.e. salient and infrequent events), whereas a study focused on primary care visits (i.e. typical and frequent events) would ideally be limited to a 6 month, or less, recall period (12, 19). In cases where the utilization events of interest are regular and predictable, such as consumption of prescription drugs, monthly utilization data could be used to infer annual utilization (11). However, this would introduce substantial estimation error with utilization activities that are irregular or subject to seasonal variation, like physician visits. In addition, if time and availability of resources permit, the level of recall bias could be quantified and the self-reported data could then be inflated or deflated accordingly. For example, Brusco and Watts (10) quantified the level of under-reporting of self-reported visits to a general practitioner by comparing with national claims data, leading to the recommendation to inflate such visits by 16%. Many of the techniques described earlier were used in a study aimed at identifying groups of lung cancer patients with similar diagnostic pathways in the primary care interval—from first presentation in primary care with signs and symptoms suggestive of lung cancer to specialist referral (31). The grouping of patients was based on health care service utilization patterns with a focus on visits to general practitioners, hospitalizations and imaging tests. A major issue in relation to recall bias in this study was the length of the recall period. Although this was set at a maximum of 1 year, the uniqueness of each patient’s pathway led to a large variation of recall timeframes within the sample. For example, some patients were referred the same day they presented in primary care, whereas others had long delays before referral. Moreover, recruited patients were already diagnosed with lung cancer meaning that the time between the utilization activities and the interview may have been up to 3 years. For patients diagnosed much earlier, this led to difficulties in grounding them in the specific recall period. Accordingly, the following measures were used to reduce recall bias. In the preparation phase, the interview guide was pilot tested and revised to ensure questions were clear and precise. This was done to reduce variation in comprehension and allowed practice and evaluation of interviewer technique to ensure consistency across interviews. Approximately 1 week before the interview, patients were asked to think about past visits and have their personal agenda or other helpful items with them at the time of the interview. This served to enhance recall and reduce the risk of under-reporting, especially in patients who were older. During the interview, patients were asked to self-report their health care utilization starting with their first presentation in primary care with signs and symptoms suspicious of lung cancer and working forward month by month to the date of specialist referral. In addition, the interviewer brought a large calendar to record the data in a collaborative manner with the patient. This allowed a visual of the causal sequence of events and facilitated event recall around landmark events such as birthdays, or routine activities. An example of this is shown in Figure 1 where the patient recalls an appointment with their family physician the day after a regularly scheduled bridge game. Calendar data illustrating event recall around a routinely scheduled activity. These measures necessitated a face-to-face interview where we were also able to include a caregiver/companion who was knowledgeable about the patient’s health care utilization during the period of interest. This helped to improve data accuracy, especially in cases where there may have been mild cognitive impairment. Self-reported health care utilization is a viable method that can be used in a range of studies; however, the potential for recall bias must be acknowledged and reduction measures to address it must be incorporated. We have presented various types and sources of recall bias and have provided measures that can be easily incorporated into a study to promote data accuracy. The example in primary care presents several challenges, such as length of the recall period, and offers concrete strategies to offset such challenges. We have summarized these points in Figure 2 as a helpful guide towards measures that can be taken when trying to reduce recall bias. Modifiable attributes as per Bhandari and Wagner (12) and specific measures to improve data accuracy using the example of Khare et al. (31). Funding: The study discussed in this article was supported by the Fonds de recherche du Québec -Santé (FRQS) and the Unité de soutien à la stratégie de recherche axée sur le patient (SRAP). Ethical approval: The study discussed in this article received approval from the Research Review Office of the Integrated Health and Social Services University Network for West-Central Montreal. Conflict of interest: The authors have no conflicts of interest to declare.
No takes yet. Share an insight, caveat, or question.
Khare et al. (2019) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: