Key points are not available for this paper at this time.
Design
Editorial
This educational piece explains the concepts of heterogeneity, fixed-effect, and random-effect models in meta-analyses.
Dear Ms. Method Matters, I saw a recent meta-analysis published in Anaesthesia on the analgesic effects of propofol total intravenous anaesthesia (TIVA) compared with inhalational anaesthetics 1. I noticed that in some of the studies included in the meta-analysis, the primary outcome, pain, was assessed using a numerical rating scale, while other studies used a visual analogue score (from 0 to 100, or 0 to 10), instead 1, 2. In addition, the meta-analysis 1 pooled data from studies which investigated the analgesic effects of propofol after different surgical procedures, some of which may have very different pain outcomes; for example, some studies were looking at laparoscopic surgery, whereas others were open surgical procedures. Can data from these studies be collated into a meta-analysis, can using fixed or random-effect models overcome this problem, or are the studies so different from each other that we would be comparing apples and oranges? Mystified with meta-analyses (Southampton) Dear Mystified, Meta-analysis is defined as the statistical analysis of a large collection of analysis results from individual studies for integrating the findings 3, 4. As the number of clinical studies with similar protocols has grown, meta-analyses have become increasingly popular, since more reliable results can be obtained by systematically combining studies and the limits of size and scope of individual studies can be overcome. However, the goal of a meta-analysis is not simply to report the mean effect size of an intervention; it is also important to know how the effect sizes in the various studies are dispersed around the mean: in other words, is the impact of the intervention consistent, or does it vary across studies? The consistency of results across studies needs be determined to ascertain whether the findings can be applied in other populations. In an attempt to establish whether the results from studies are consistent, meta-analyses will often include a measure of heterogeneity, for example, Cochran's Q, or I2 5. The test statistics Cochran's Q or I2 determine whether there are true differences underlying the results of the collated studies (denoting heterogeneity), or whether the differences observed between studies are due to chance or sampling error (denoting homogeneity) 5. Tests for heterogeneity are tests for the null hypothesis that all studies are investigating the same effect. Since meta-analyses pool data from studies which are diverse methodologically as well as clinically, some degree of heterogeneity is to be expected. The classical measure of heterogeneity is Cochran's Q, which is calculated as the weighted sum of squared differences between individual study effects and the pooled effect across studies. Cochran's Q has low power as a test of heterogeneity 6 when the number of studies is small, but too much power if the number of studies is large 5. Since the reliability of Cochran's Q as a measure of heterogeneity depends on the number of studies included in the meta-analysis, other measures of heterogeneity have been developed, the most popular of which is I2, (I² = 100% × (Q-degrees of freedom)/Q. I² is presented as a percentage, and the interpretation is simple and intuitive. A rough guide for the interpretation of I2 is: 0–40%, may not represent clinically important heterogeneity; 30–60%, is considered moderate heterogeneity; 50–90%, may represent substantial heterogeneity; and 75–100% is considerable heterogeneity. Most meta-analysis software will include calculations of Cochran's Q, or I2, and it is always a good idea to look at the values given before deciding on how to proceed with the meta-analysis. Should substantial heterogeneity be found, authors should revisit the primary studies to establish what is causing the heterogeneity and decide whether that particular study should be included in the meta-analysis. Researchers should note that it is possible to obtain an I2 score of 0% because of the way the score is calculated, and although there is not necessarily anything amiss, they should nevertheless revisit the primary studies to ascertain whether an I2 score of 0% is feasible. If heterogeneity is low, a researcher can decide to proceed with data analysis using a ‘fixed-effect’ model. This model presumes that the studies under investigation were all conducted under similar conditions, with similar subjects, which means that the only difference between studies is their power to detect the outcome of interest. In practice, it is very difficult to find so many published studies which have used similar protocols and similar subjects, using exactly the same primary outcomes, and unless the studies were conducted on genetically similar animals, some heterogeneity is inevitable. An alternative approach to the ‘fixed-effect’ model is the ‘random-effect’ model. The random-effect model allows the study outcomes to vary in a normal distribution between studies and its use can be considered if heterogeneity is high 7. A goal in meta-analysis is to estimate the combined effect, and if all studies included in the meta-analysis were equally precise, the combined effect of all the studies can simply be reported. However, should the studies be heterogeneous, more weight should be assigned to studies which are more precise 7. In a fixed-effect model, because it is assumed that there is one true effect size across all the included studies, the combined effects are the estimate of the common effect size and the weights assigned to each study will depend solely on the sample size of the study. A study with more subjects will be assigned a higher weight, and a lower weight will be assigned to a study with fewer subjects 7. In a fixed-effect model, large studies are very likely to dominate the results of a meta-analysis. It is presumed that, in a fixed-effect model, the only source of error is the random error within the studies, and this error will tend towards zero as the size of the study becomes bigger, and this is true no matter whether the ‘size’ of the study (i.e., the meta-analysis) is contributed to by one, or many individual primary studies. However, in the random-effect model, because it is assumed that the effects will vary according to a normal distribution, the studies are a random sample of the relevant distribution of effects. In contrast to a fixed-effects model, large studies may yield a more precise estimate than small studies, but each study is actually estimating a different effect size. Each study, no matter how small or large, serves as a sample from the population whose mean is being investigated. In the random-effect model, there are two levels of error: the first level being that each individual study is used to estimate the true effect in a specific population; and the second level of error being the estimate of the mean of the true effects across studies 7. Generally, it is implausible to assume in clinical studies that the true effect is the same in all studies, and more often than not, a random-effect model is used 8, 9. In the example as quoted above 1, post-surgical pain was compared in patients undergoing surgery anaesthestised using either inhalational anaesthetics or propofol, but the type of surgical procedures investigated as well as patient populations, were not identical. In such a scenario, should there have been enough studies under each surgical category 10, a sub-group analysis could have been undertaken, however, the risk of bias also needs to be taken into account 11. For example, all the studies investigating pain after orthopaedic surgery could have been analysed together, all studies investigating pain after mastectomy could have been analysed together and so on (as in Lam et al. 12). It is to be noted though that calculations for heterogeneity should still be conducted for each sub-group (i.e. an I2 value should be reported), and decisions on random, or fixed-effect model must be made based on the individual sub-group heterogeneity scores 7. It is true that levels of pain encountered after different surgical procedures vary substantially 13, but, if a random-effects model is used, assuming that the analgesic effect of propofol TIVA would follow a normal distribution, collating these individual studies in a meta-analysis would not be akin to comparing apples and oranges. If you have any questions on methodology, please direct them to Ms Method Matters at msmethodmatters@gmail.com. SC is statistical advisor to Anaesthesia. No other competing interests declared.
No takes yet. Share an insight, caveat, or question.
Choi et al. (2017) studied this question.