The insightful and stimulating commentary by Julian Higgins 1 on our paper 2 raises several important issues that need to be clarified. First, we need to agree on nomenclature. The heterogeneity literature has been plagued by inconsistent terminology. Terms like heterogeneity, inconsistency, variation, diversity, between-study variance, variability, etc. are used interchangeably. While Higgins prefers the term inconsistency for I 2, in other writings he has used the words variability and heterogeneity in association with this measure. 3 We believe that the term heterogeneity is a nice word with roots going back to ETEPOGENHS of Aristotle and ETEPOGENOS of Sextus Empiricus. It can be applied to any of the popular metrics and tests, but then one simply has to specify which metric or test is exactly alluded to. Inconsistency is also a nice, more recent word, but again we need to clarify what it refers to each time. Higgins worries about the post hoc hypotheses that need to be thought up to explain why the excluded studies might be outlying or influential. We were clear cut in our paper that this is indeed not an easy task. We believe that sensitivity analyses, as currently performed, are usually an invitation to post hoc data dredging with few or no rules in the game. This reduces their inferential reliability. However, this is a major reason why our proposed algorithms may offer one way to improve this free-lunch situation. There are two components to any sensitivity analysis. The first component is how it is done. The second component is how the results are interpreted. ...
No takes yet. Share an insight, caveat, or question.
Patsopoulos et al. (2008) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: