I believe that we do not know anything for certain, but everything probably. ―Christiaan Huygens, Dutch Physicist The statement “trends toward significance” needs to be put to rest. We put so much weight on that 0.05 chance, but is that the right thing to do? The p value is just a probability of the effect found in a study.1 In simple terms, it is an indication of how incompatible the data are with the null hypothesis. A small p value indicates that the null hypothesis is unlikely given the data collected.2 A 95 percent chance of a real difference is a fine goal to aspire to; however, a 90 percent chance is not something to be taken lightly, depending on the data being presented. In randomized controlled trials that have large mean differences, low-risk interventions, and life-threatening implications, surgeons should take data seriously even with p values as high as 0.1. An example of this would be a randomized trial evaluating the efficacy of mechanical prophylaxis on venous thromboembolism risk. In contrast, a meta-analysis with large numbers of patients, high-risk or expensive interventions, and minor treatment implications, a low p value is an important qualification for a lower alpha. An example for this would be a study evaluating highly specialized and invasive imaging and its effect on decreasing free flap harvest time by 2 to 5 minutes (Table 1). In other words, one can always present a low p value if the sample size is large enough, such as in a national population study using administrative data. Even if the effect size or the difference between treatments is small, the p value may be significant at p < 0.05. Table 1. - Factors Affecting When a High versus Low Significance Level Is Most Appropriate Significance Level 0.05 0.1 Meta-analysis X Randomized controlled trials X Intervention is high risk X High treatment implication X Low event occurrence X It can be difficult for plastic surgeons to obtain a high enough sample size to adequately power a study that can obtain a p value of 0.05. The commonly used alpha values of 0.05 and 0.01 are based solely on tradition from other scientific disciplines. The p value is affected by the sample size and the effect size, and both need to be deeply understood when evaluating what p value is relevant.3 Effect Size and Confidence Intervals The effect size of a study is the magnitude of the observed phenomenon. A reduction in infection rate from 1.5 percent to 1.4 percent with antibiotics translates to a low effect size, regardless of how one measures it. This may become statistically significant with a large enough sample size. However, it says nothing about the clinical relevance, or the risk-to-benefit ratio of providing antibiotics. This decision needs to be made clinically by the physician. One should not be mistaken by a p value of 0.00001 to denote a large effect; it simply notes the probability of the effect, and if the effect is small based on clinical observation, no matter how small the p value is, the effect is still clinically insignificant. Recent studies have evaluated the data in the plastic surgery literature and made smart proposals on where we can continue to improve.3–5 One important suggestion is to mandate the reporting of confidence intervals so that we can better understand the range of doubt in effect sizes and help clinicians determine whether the intervention is practically meaningful. In the top seven plastic surgery journals, reporting of confidence intervals is only 25 percent.3 Once we give better data and can understand effect size, we should be able to increase relevant p values for certain study types as described above. What is the right significance level? The most difficult interventions to study are those that evaluate events with low occurrences in operations that are not performed commonly. For example, free flap losses occur at approximately 1 to 5 percent in high-output centers; thus, adequately powering a study to correlate age with free flap losses could only be done by one or two large institutions and is unlikely to ever be replicated. This could at least partially explain the inconsistency in reporting complications in the plastic surgery literature.4 Let us say that you have an intervention that could reliably reduce your free flap losses from 5 percent down to 0.5 percent. To power a two-group study to 80 percent with an alpha of 0.05, you would need to have a very large sample size of 412. Changing your alpha to 0.1 brings the sample size to a still large but more manageable 324 patients. Most interventions for these rare event occurrences do not have this profound of an effect. The p value should be raised to 0.1 for studies that fit this mold. In these cases, if the results support the conclusion, the authors would suggest that the data are significant to a value of 0.1 rather than stating that the data “trend toward significance.” The most obvious disadvantage of raising the p value is the real increase in type I error. More false-positive results will lead to more poor data being submitted to journals. The onus will be on reviewers and editors to determine whether the greater significance value is adequate. This should be mitigated by allowing meta-analysis studies to pool the wider body of published research. With more data in the published literature, better meta-analysis studies can be performed with more pooled data.5 CONCLUSIONS Changing a plastic surgeon’s preconceived notions about a procedure he or she has been performing his or her whole career is the most difficult task in scientific writing. Even if we do not change the significant p value to 0.1, we need to do a better job of interpreting and publishing negative results. The data should leave it up to the readers to determine whether the 90 percent chance of a real difference is enough to change their practice for the specific question at hand. The resulting p value must be weighed against the study’s question and design, which requires an understanding of the data collection, management, and test being performed.2 In other words, there is no one metric such as the p value to guide the validity of the study. The reviewers must be avid consumers of study designs, with their inherent strengths and biases. Only when one understands how a study is conducted can one deduce how accurate the conclusions are. As reviewers and editors, we can help ensure that data are interpreted more accurately. The first step would be to determine which studies would be better understood with a statistically significant value of 0.1 instead of 0.05.
No takes yet. Share an insight, caveat, or question.
A 2020 study studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: