The number needed to treat (NNT) is a method of reporting outcomes from clinical trials.1 Treatment efficacy is determined by evaluating the outcome of one treatment relative to another treatment or to a control group when the only difference between the groups is the intervention of interest. The NNT can also be used to express the size of the outcome of one treatment relative to another. The NNT is expressed in terms designed to help decide whether the intervention might be valuable in clinical practice: the number of patients who need to be treated before a therapist can be sure that one patient improved who would not have improved without the intervention. For example, when comparing treatment X and treatment Y, an NNT score of 5 for treatment X indicates that, on average, after treating 5 patients, treatment X will have achieved one more positive outcome than if treatment Y had been used. The NNT does not tell the clinician which of those 5 patients will respond, only that 1 patient is likely to do so. The NNT was described in 1988 by Laupacis et al,2 and, although its use is becoming more popular, it is still not widely used. A search of MEDLINE back to 1991 using the search terms “number needed to treat” or “NNT” identified 121 citations reporting NNT information. Of these citations, only 3 concerned outcomes of physical therapy3–5 and, of these, only one5 described the use of NNT in a journal with content relevant to physical therapists. The NNT can be used to gauge the relative effectiveness of different treatments in restoring normal function or preventing future disability. If the treatment has a potentially harmful outcome (eg, stroke, headache, death from cervical manipulation), a similar calculation can be used to indicate the number of patients who need to be treated in order to harm or even lead to the death of one person. If the statistic is used in this way, it is referred to as the number needed to harm (NNH).2,6 Treatment effectiveness can be effectively studied through the use of randomized control trials (RCTs). When the outcomes of an intervention are studied using an RCT, the various characteristics—both known and unknown—of the control and intervention groups are randomly distributed, except for the intervention of interest, which is applied to only one group. When the number of positive outcomes in each group is determined, the RCT design allows for the positive outcome of the intervention of interest to be calculated by subtracting the outcome in the control group from the outcome in the intervention group. The two halves of the denominator indicate the proportion of positive outcomes in each group individually. When the proportion of successful outcomes in the control group (or alternative therapy group) is subtracted from the proportion of successful outcomes in the intervention group, it indicates the relative efficacy of the intervention. This is the outcome that can be attributed solely to the treatment under investigation. If the control group performed better than the treatment group, a negative value is produced, indicating that the treatment may be ineffective or harmful. The denominator, therefore, indicates the attributable outcome of the treatment. To render this value applicable in the clinical setting, the number of subjects who must be treated before one extra subject will be helped by treatment is calculated by dividing the denominator into a numerator of 1. Thus, NNT is the inverse of the outcome attributable to the treatment alone. If 100% of the subjects in the intervention group respond positively, whereas none do so in the control group, the NNT=1/1=1, indicating that every patient treated responds favorably to the intervention. As the difference between the positive outcome due to the intervention and the positive outcome from the control group decreases, the NNT increases. We will use what we consider a well-designed RCT by Watt et al8 to demonstrate calculation of the NNT. One aim of their study was to compare the outcome of physical therapy intervention with that of a home exercise program on wrist extension range of motion (ROM) following Colles' fracture. After removal of the plaster cast, a measurement of wrist ROM was done. Subjects were then randomly assigned to either a physical therapy group (n=9) or a control group (n=9) using sealed envelopes that concealed the random assignment. A computer-generated random number list was used to order the envelopes. The orthopedic surgeon or registrar gave the control group a written exercise program to perform at home. The physical therapy group received treatment at the discretion of the treating physical therapist. Treatment typically included active exercises in a home program, home advice, and passive joint mobilization. The control group did not receive any individualized physical therapy. The amount of exercise performed by each group did not differ between groups (t(16)=1.63, P=.12). For all subjects, wrist extension ROM was measured with a goniometer by one investigator who was unaware of the group assignment for each subject. After 6 weeks, wrist extension ROM was measured again. Table 1 shows the results of wrist extension ROM measurements from the study by Watt et al.8 Before intervention, the mean wrist extension ROM was 30 degrees and 28 degrees in the physical therapy and control groups, respectively. At the 6-week follow-up measurement, the mean wrist extension ROM was 55.7 degrees and 38.3 degrees in the physical therapy and control groups, respectively. Wrist Extension Range of Motion (ROM) Measurements From Watt et al8 Wrist Extension Range of Motion (ROM) Measurements From Watt et al8 Watt et al8 analyzed the data using a one-tailed split-plot analysis of variance (SPANOVA). A positive outcome of physical therapy was reported (F=2.59; df=1,16; P=.010) (Tab. 1). The immediate implication of this significant result for clinical practice is not obvious. The numbers reported (F and P values) do not translate into estimates of the magnitude of the positive outcome of the intervention. The reported results do not indicate the size of the positive outcome or the percentage of patients who would benefit from physical therapy in addition to exercise intervention. The reported results only indicate that an outcome in favor of physical therapy has occurred. Kapandji9 suggested that prehension is optimized with the wrist in 40 to 45 degrees of extension. Therefore, if a wrist ROM of greater than 45 degrees of extension is defined as a positive outcome for the study by Watt et al,8 then NNT can be calculated with the formula using the information in Tables 1 and 2. Summary of Data From Watt et al8 Required for Calculation of Number Needed to Treat (NNT) Summary of Data From Watt et al8 Required for Calculation of Number Needed to Treat (NNT) This calculation indicates an NNT of 2.25, meaning that clinicians need to treat 9 patients as described by Watt et al8 before they can be confident of achieving 4 positive outcomes of wrist extension greater than 45 degrees that would not normally be achieved with home exercises. Compared with the result of the one-tailed SPANOVA, the NNT provides results in a way that is directly transferable to the clinical setting. Although it is useful to know whether a treatment is effective, knowing how useful the treatment is in terms of the numbers of patients who need to be treated helps in calculating how much the expected gains will cost, both in financial terms and in terms of the demands on the patient to continue the therapy. This is the great advantage of the NNT. If a person wanted to change the definition of a positive outcome (eg, doubling wrist extension ROM in 6 weeks), it is possible to recalculate the NNT. Table 3 illustrates how different definitions of a positive outcome yield different NNT scores from the same data. Of interest is the result of the last 2 positive outcomes in the table. When comparing the outcomes, achieving wrist extension ROM greater than or equal to 30 degrees would seem an easier goal to achieve than increasing ROM by greater than 50 degrees. Yet, achieving wrist extension ROM greater than or equal to 30 degrees is associated with an NNT of 3, whereas achieving 50 degrees or more of ROM is associated with an NNT of 1.5. The NNT of 3 tells us that more patients in the control group will achieve the positive outcome of 30 degrees and, therefore, more patients need to be treated to find that one extra patient who would not have achieved the outcome ROM by doing exercises alone. When 50 degrees of ROM is the goal, however, the benefits of the physical therapy intervention are more apparent. Here, it is more unlikely for a subject from the control group to achieve this goal, but a high proportion of subjects will achieve this goal in 6 weeks if they have physical therapy. Contrast these results with the positive outcome defined as doubling ROM. The entire sample would need to be treated by a physical therapist before one extra person would double his or her ROM. Examples of How Different Definitions of Positive Outcomes Yield Different Numbers Needed to Treat (NNT) Using the Data From Watt et al8 Examples of How Different Definitions of Positive Outcomes Yield Different Numbers Needed to Treat (NNT) Using the Data From Watt et al8 Once the SE is known, calculation of the upper and lower confidence limits using equation 4 is possible.11 Using this process, CIs for NNT for the study by Watt et al8 can be calculated. The upper NNT limit is 16.6, and the lower NNT limit is 1.21. These confidence limits can be interpreted to mean that the true NNT might be as high as 17 or as low as 2. Because of the inverse nature of NNT, the CIs are not symmetrical around the point estimate of 2.25. Physical therapists are often interested in the preventive value of their interventions, and NNT can also be used to interpret these results. The utility of interventions designed for prevention is reflected in the rate or number of non-events. The formula for calculating NNT remains the same. The use of NNT in this manner is illustrated by a study that investigated whether increasing flexibility can reduce the rate of lower limb overuse injuries in a population of army recruits.12 Two different army companies that were completing basic infantry training at the same time were studied. The control group completed normal training, whereas the treatment group had 3 extra stretching sessions incorporated into their weekly training program. The length of basic training was 13 weeks. A physician who was unaware of group allocations, recorded the number of injuries presumably due to overuse in both groups during the 13 weeks of training. Overuse injuries were defined as injuries consisting of stress fractures, patellofemoral pain, muscle strains, tendinitis, plantar fasciitis, shin splints, and anterior compartment syndrome. The results of the study are summarized in Table 4. Results of Study of Overuse Injuries12 Number of positive outcomes=group size−number of injuries. Results of Study of Overuse Injuries12 Number of positive outcomes=group size−number of injuries. An NNT of 8 indicates that for every 8 trainees undergoing the preventive stretching regimen, 1 trainee was prevented from an overuse injury of the lower limb during the 13-week period. The results indicate the potential benefit in changing basic infantry training.12 The cost to rehabilitate a recruit after sustaining an overuse injury would be more than the cost of allowing 8 recruits to perform stretching exercises 3 times a week during basic training. Confidence intervals can also be calculated for this study in the same manner described above. In this study, the upper NNT limit is 33.9 recruits, and the lower NNT limit is 4.6 recruits. In this sample, 43 recruits from the control group had overuse injuries. Therefore, the risk of sustaining an overuse injury in this sample is 0.29. An NNT score of 8, as calculated for the data from Hartig and Henderson,12 informs the reader of the outcome of 1 of the 8 recruits. It does not indicate the response to intervention of the other 7 subjects. Of the 7 recruits, some are likely to develop an “overuse injury,” given the risk of this injury in basic training. Thus, the NNT reflects the average number of subjects who need intervention to prevent one event, but the NNT does not inform the reader of the fate of the other subjects in the sample. Because the subjects who will sustain the adverse outcome cannot be predicted in advance, all 8 subjects would need to have interventions in order to prevent one adverse outcome. An NNT score of 1 means that every time a treatment is used on the defined patient group, it will result in a desired positive outcome that would not have occurred without treatment. Treatments with NNT scores approaching 1 are found when comparing antibiotics with placebo in the treatment of Helicobacter pylori infection.6 Some authors believe that an NNT score of 3 or less indicates a worthwhile intervention.6 What is considered a worthwhile intervention depends on the definition of a positive outcome. If an intervention with an NNT of 10 can save one extra life from a disease that is killing millions, then the intervention would seem to be worthwhile, despite the NNT score being larger than 3. The reader needs to weigh the relative merits of the intervention and its possible outcomes when interpreting NNT scores. As with the results of other statistical tests, the observed outcome must also be considered, but NNT provides information that offers clinically meaningful insights. If treatments are shown to have NNT scores lower than 3, then it may be worthwhile to consider instituting a change in clinical practice if other factors do not indicate otherwise.6 However, the decision to change clinical practice must be weighed against the potential harm—expressed as the NNH—and the costs of this change and how these factors could affect health care resource allocation in general. As the NNT score increases, more resources are required to obtain one new positive outcome in the defined time period. When treatments are shown to be weakly effective, ethical and financial considerations also need to be taken into account, and here a treatment with little resource needs might seem more attractive even without a good NNT. The NNT information aids decisions regarding appropriate clinical practices and optimum utilization of available resources by expressing the results from trials in terms of patients needed to be treated. The constraints of a study also dictate how NNT should be interpreted. In the study by Watt et al,8 only outcomes obtained over a 6-week period were considered. Therefore, utilization of NNT information depends on the quality and design of the research. The NNT provides clinically relevant information only when clinically relevant studies are performed. The NNT is a method of reporting and interpreting outcomes of clinical trials and, in our opinion, provides information in an easily understandable form. The NNT can be used to describe results from studies that explore the preventive outcome of treatment and those that explore treatments designed to restore normal function.
No takes yet. Share an insight, caveat, or question.
Dalton et al. (2000) studied this question.