Researchers make hundreds of decisions about data collection, preparation, and analysis in their research. We use a many‐analysts approach to measure the extent and impact of these decisions. Two published causal empirical results are replicated by seven replicators each. We find large differences in data preparation and analysis decisions, many of which would not likely be reported in a publication. No two replicators reported the same sample size. Statistical significance varied across replications, and for one of the studies the effect's sign varied as well. The standard deviation of estimates across replications was 3–4 times the mean reported standard error.
No takes yet. Share an insight, caveat, or question.
Huntington‐Klein et al. (2021) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: