Looking through any exercise science journals today, in fact any science journals including many top Science Citation Index (SCI) journals, one can easily find examples of the wide-spread "p < 0.05/significance" abuse phenomenon, i.e., if the p value from a statistical/hypothesis test is less than 0.05 (or 0.01 sometimes), a conclusion that "the results/ findings are significant" is then drawn.The abuse is so severe that it is already seriously threatening the integrity of scientific inquiry.Why is the popular p value practice a problem?An example may help to explain.When I teach my graduate research methods class, I usually conduct a survey about students' background on my first day's class so that I can prepare my teaching according to the students' background and needs.Two of the questions in the survey are about the students' undergraduate Grade Point Average (GPA) and the Graduate Record Examinations (GRE) scores.Table 1 illustrates 14 students' responses in 1 year's survey.Say if I am interested in knowing the impact of undergraduate training on students' GRE test performance, I can run a correlation between GPA and GRE using the data in Table 1.The correlation coefficient (r) is 0.178, with a p value of 0.544.Since the p value is larger than 0.05, we can then conclude that there is no relationship between GPA and GRE.But let's go further and do a small experiment: We simply copy the sample data and paste them into the existing data set to increase the n in the statistical software we are using, and re-compute r and p value each time (Note: This experiment is only trying to make my point and SHOULD not be done in a real study!).We repeated this process eight times and summarized our computational results in Table 2.
No takes yet. Share an insight, caveat, or question.
Weimo Zhu (2012) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: