This article implies that sharp inferences to large populations from small experiments are difficult even with probability sampling. Features of random samples should be kept in mind when evaluating the extent to which results from experiments conducted on nonrandom samples might generalize.
No takes yet. Share an insight, caveat, or question.
Tipton et al. (2016) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: