Key points are not available for this paper at this time.
Introduction The efficacy of HAART has improved over the past 10 years, with the introduction of more potent drugs with improved safety profiles and lower pill counts 1. In the most recent clinical trials of naive patients, 48 week HIV RNA suppression rates below 50 copies/ml of over 80% have been recorded 2,3. Similar improvements in viral suppression rates have been reported from clinic cohorts 4. Where combinations of novel agents have been used, short-term efficacy rates for treatment-experienced patients have now risen to levels similar to those seen in naive patients 5. Studies of treatment naive patients suggest an upper limit of virological efficacy, beyond which the inclusion of new or additional drugs may provide only small incremental benefits. For example in the ACTG 5095 trial, there was no additional benefit for naive patients given zidovudine/abacavir/lamivudine/efavirenz versus the control arm of zidovudine/lamivudine/efavirenz 6. Given these developments, any superiority trial designed to demonstrate improved efficacy with a new drug combination would have to be large. Consequently the ‘noninferiority’ trial is emerging as a new standard design for HIV drug development among antiretroviral-naive individuals. These trials are designed to demonstrate that a new treatment shows efficacy not substantially worse than the current standard, within a prespecified margin (also called the noninferiority margin, or ‘delta’). The new treatment may offer other benefits, such as simplified dosing, an improved safety profile or a lower incidence of drug resistance at treatment failure 7,8. This outcome is best illustrated with an example from the EPV20001 trial (results shown in Table 1), in which efficacy rates of 64% for lamivudine once daily and 63% for lamivudine twice daily (both in combination with zidovudine and efavirenz) were observed. This result is expressed as a point estimate (a 1% benefit for the lamivudine once daily arm) and a 95% confidence intervals (between a 7.1% disadvantage and an 8.9% advantage for the lamivudine once daily arm). In this case, the lower 95% confidence limit of a 7.1% disadvantage for the once-daily arm is less than the 12% delta set in the original protocol, so lamivudine once-daily was declared noninferior to the then standard treatment of lamivudine twice daily (Table 1). The benefit of lamivudine once daily was then more convenient dosing with no substantial loss of efficacy compared with the control arm of twice daily lamivudine dosing. This outcome is also shown as scenario C in Fig. 1.Table 1: Summary of phase 3 company-sponsored noninferiority trials (2000–2007): statistical design and 48-week efficacy intent-to-treat (ITT) – time to loss of virological response (TLOVR) analysis.Fig. 1: Examples of the four main outcomes from noninferiority trials.There is, however, much inconsistency in the design and interpretation of HIV noninferiority trials 7. In this review, we describe the design of these studies and their interpretation, and discuss the implications of this design for the choice of endpoints and sample size calculations. Our aim is both to educate and review the use of such trials in the HIV setting. Review methods The present review includes data from company-sponsored Phase 3 noninferiority trials conducted between 2000 and 2007, defined as Phase III/IV trials conducted with company sponsorship for both trial conduct and study drugs. Only the company-sponsored trials were included as they normally follow US Food and Drug Administration (FDA) guidelines on design and reporting of HIV RNA endpoints 9, and could therefore be interpreted in a standardized way. We used a MEDLINE search with the search terms of each antiretroviral, followed by ‘clinical trial’ (e.g. ‘lamivudine clinical trial’). In addition, we searched the FDA product labels for registrational trials of each approved antiretroviral, and searched for abstracts on clinical trials presented at the following conferences: Annual Conference on Retroviruses and Opportunistic Infections, International Conference on Antimicrobial Agents and Chemotherapy (ICAAC), European AIDS Clinical Society, International AIDS Conference (including IAS Pathogenesis Conference) and International Conference on Drug Therapy in HIV Infection. The search identified 17 randomized trials with a noninferiority design that used an endpoint of HIV RNA suppression below either 400 or 50 copies/ml (the design of these trials and summary efficacy data are shown in Table 1). In two cases 10,11, the trials were powered to show equivalence but can be interpreted in terms of noninferiority. All but three of the trials 11–13 were conducted in treatment naive patients. The BMS-034 trial (atazanavir versus efavirenz in naive patients) was excluded from this analysis owing to problems in validation of the HIV RNA assays used 14. There were four additional trials which were powered on noninferiority but used continuous log reduction as the primary endpoint 15–18. The design of noninferiority trials Standard superiority trials are designed to ensure that the smallest true difference between the new and standard treatments thought to be clinically relevant has a high chance of being detected as statistically significant. For clinical trials of antiretrovirals, the most commonly used endpoint is HIV RNA suppression below 50 copies/ml without treatment discontinuation. Decisions about the likely efficacy of the new treatment are generally based on a standard hypothesis test and the resulting p-value (with confidence intervals provided to aid the clinical interpretation of the findings). In contrast, noninferiority trials are designed to show that a new treatment is not substantially inferior to the current standard. Instead of focussing on the results of a statistical test, emphasis is placed on ensuring that the lower limit of the confidence interval for the observed difference in outcomes between the two regimens does not cross the prespecified ‘delta’ 19. Table 2 shows sample sizes required to show noninferiority for an experimental treatment versus control, given percentage response rates in the control arm ranging from 50 to 90%, and delta of either 10 or 12%, consistent with current FDA and European guidelines 8,9,19. These calculations assume that the true response rate is equal in the experimental and control groups. In some trials, the response rate is assumed to be slightly lower in the experimental arm, which leads to larger sample sizes – this design was included in the KLEAN trial 20.Table 2: Sample sizes per arm for noninferiority trials, by power, delta and expected response rate in the control arm; the efficacy of the new drug is assumed to be equivalent for the purposes of calculating sample sizes.The choice of an appropriate delta may be problematic: delta is usually chosen to reflect the largest difference in outcomes between the arms that could reasonably be assumed to be clinically equivalent 19. A few trials have been designed with a high delta – for example the BI 118.33 trial was designed to show that tipranavir/ritonavir was no more than 15% worse than lopinavir/ritonavir 21; the design of the Abbott 418 trial of once versus twice daily lopinavir/ritonavir also included a delta of 15% 22 (Table 1). Trials powered with a delta of this size may not, however, be able to exclude the possibility that a true difference exists between the arms which may be considered to be clinically significant. For example, if a noninferiority trial showed that the efficacy of first-line tenofovir/emtricitabine/efavirenz was 80% but that the true efficacy of a new combination treatment may be as low as 65%, would this convince clinicians to choose the new combination over the current standard, even if it were deemed to be noninferior? In addition, the noninferiority margin (delta) should be smaller if the control arm is already highly efficacious (i.e. with response rates above 90%). In future, differences between treatment arms of less than 10% may be considered clinically significant, and therefore a delta of 10–12% could be considered too large. For example in the Gilead 934 trial, there was a 7% advantage in efficacy of tenofovir/emtricitabine/efavirenz over the control arm of zidovudine/lamivudine/efavirenz at week 48 2, driven mainly by higher rates of anaemia and gastrointestinal toxicties in the zidovudine arm. As a result, zidovudine is no longer recommended for first-line use in Europe 23. Once the delta falls below 10%, however, Phase III trial sample sizes could rise to levels where the economics of HIV drug development become unsustainable. For example if a new experimental drug was compared to tenofovir/emtricitabine/efavirenz with a predicted success rate of 80%, and a delta of 5%, the trial sample size would be 1005 patients per arm for a power of 80%, and 1345 patients per arm for a power of 90%. Intent-to-treat and per protocol analysis When performing a superiority trial, the primary analysis uses the intent-to-treat population, including all patients randomized, irrespective of whether they have taken their study medication as randomized. Such an approach tends to bias the results towards the null hypothesis (i.e. no difference in outcome between the treatment arms). Thus, if there is still a difference in outcome when the trial is analysed in this way, it is likely that the real difference (if all patients were able to take the drugs as planned) would be greater. Unlike a superiority trial, noninferiority trials usually favour a ‘per protocol’ analysis. This analysis excludes patients with major protocol violations, such as not receiving at least one dose of study drug or using a disallowed medication in the background regimen 8. By excluding these patients (who would be expected to make the two groups more alike), it is thought that analysis of the per protocol population may be more likely to show differences between treatments. For noninferiority trials, demonstration that the new treatment is noninferior on both the intention-to-treat and per protocol populations is usually required. There is, however, no standard predefined list of the exclusions for a per protocol analysis. Some per protocol analyses only exclude patients with the strongest protocol violations, such as not taking one dose of randomized treatment 12,24, whereas other analyses exclude from the analysis all individuals experiencing a nonvirological endpoint 13. In addition, in order for the results from the per protocol analysis to be of value, it is important that a large proportion of the patients randomized in the trial fall into the per protocol population - demonstration of noninferiority on only a small minority of randomized patients is unlikely to convince many clinicians that the new regimen is genuinely noninferior. Furthermore, poorly conducted trials also tend to obtain results that are biased towards ‘no effect’. Therefore, both careful support for patient adherence and attendance, including documentation to support this, and evidence of strict adherence to the study protocol, are of utmost importance in a noninferiority trial. The choice of endpoints for noninferiority trials HIV RNA suppression <50 copies/ml has been adopted as the primary objective of antiretroviral treatment in the most recent guidelines reflecting the sensitivity of most currently available commercial assays. The FDA TLOVR (time to loss of virological response) algorithm has been used to analyse HIV RNA data from registrational trials 9. This algorithm classifies patients either as virological successes while taking randomized treatment (with HIV RNA below the detection limit on two consecutive study visits around the 48 week timepoint), or treatment failure, divided into three categories: Virological failure – either failure to suppress HIV RNA, or virological rebound after initial suppression. Discontinuation of randomized study treatment for adverse events or death. Discontinuation of randomized study treatment for other reasons (for example withdrawal of consent or loss to follow-up). When using composite endpoints such as this, it is assumed that all components of the endpoint are viewed as being equally detrimental – it can be argued that discontinuation of treatment due to withdrawal of consent may have less clinical relevance for future virological suppression than a virological failure. One particular limitation with composite endpoints such as these is that it can be difficult to interpret intent-to-treat analyses where the virological and nonvirological endpoints are imbalanced across treatment arms. For example in the MERIT trial the experimental treatment of zidovudine/lamivudine/maraviroc showed an excess of virological failure endpoints, whereas the zidovudine/lamivudine/efavirenz arm showed an excess of discontinuations for adverse events 24. In a recent survey, only 27% of endpoints in trials of naive patients were due to virological failure, with the remaining 73% being due to discontinuation of study medication (32% for adverse events and 41% for loss to follow up) 25. Given that trial outcomes can be dominated by nonvirological endpoints when analysed by the FDA TLOVR algorithm, it is important for HIV clinical trials to be re-analysed including only virological endpoints. The FDA guidelines state that, in addition to analyses using the TLOVR algorithm, an analysis comparing only the documented virological failures should be presented and any inconsistencies between the different analyses should be explored 9. This is also called a ‘nonvirological failures censored’ analysis. Data from patients is censored after discontinuation for reasons other than virological failure. Interpreting the results of noninferiority trials Figure 1 shows the four most common outcomes of a noninferiority trial. This figure shows the difference in efficacy between the experimental and control arms, with 95% confidence intervals. Examples of these results for HIV trials are also shown in Table 1. Inferiority shown despite a noninferiority design (scenario A in Fig. 1). One example where this outcome was seen was the CONTEXT trial of fosamprenavir/ritonavir versus lopinavir/ritonavir 15. The trial had been powered (using an endpoint of continuous log reduction in HIV RNA) to demonstrate noninferiority of the fosamprenavir arm, the 95% confidence interval for the true difference between the two arms did not overlap zero, suggesting that lopinavir actually out-performed fosamprenavir; a formal statistical test confirmed the inferiority of fosamprenavir in this trial. In situations such as these, the same statistical procedures as for unexpected superiority can be used. Failure to show noninferiority (scenario B). This result occurs when the 95% confidence interval for the true difference between the treatment arms overlaps zero, but the lower limit of this confidence interval falls below the prespecified delta value. This result was seen for tenofovir versus stavudine in the Gilead 903 trial 26, nevirapine versus efavirenz in the 2NN trial 27, and maraviroc versus efavirenz in the MERIT trial (which was reported with a one-sided confidence interval) 24. A common misinterpretation of the results from these trials is that, since the confidence interval for the difference overlaps zero, the arms are showing similar efficacy. The correct interpretation is that the new treatment failed to show noninferiority, and so the new treatment should not be accepted as an alternative to the current standard of care. Demonstration of noninferiority (scenario C). This result occurs when the 95% confidence interval for the true difference in outcomes between treatment groups overlaps zero, but the lower limit of this confidence interval remains above the predefined delta value (10 or 12%). Note that the observed efficacy of the experimental arm needs to be very close to that of the control arm (typically within 2–3%) for noninferiority to be demonstrated with this level of power and noninferiority margin. Also, a new treatment may not be accepted even if noninferiority is shown. For example in the BI 1182.33 trial, tipranavir/ritonavir 500/200 mg twice daily was shown to be noninferior to lopinavir/ritonavir, but the tipranavir arm was closed down by the Data Safety Monitoring Board owing to excess elevations in liver enzymes 21. Superiority shown despite a noninferiority design (scenario D). This result has been seen for the Gilead 934 trial of tenofovir/emtricitabine/efavirenz versus zidovudine/lamivudine/efavirenz 2, and also in the TITAN trial of darunavir/ritonavir versus lopinavir/ritonavir 12. In each case, the trials had been designed to show noninferiority, but the 95% confidence intervals for the true difference between arms did not overlap zero, suggesting superiority. The European guidelines on noninferiority trials state that it is acceptable to calculate the with a test of superiority within a noninferiority trial, provided that the superiority is not a of a lower rate of discontinuations for adverse events in the experimental arm is important that the test of superiority be on the intention-to-treat For cases these, the treatment arms are compared for noninferiority, and then for in two statistical outcome is if a treatment is shown to be worse than the control, but the confidence intervals of the difference fall within the for noninferiority not or the noninferiority We are not of an example of this in Phase III HIV efficacy trials, this outcome has been seen in studies The outcomes above may when a trial has been designed to show noninferiority. A outcome of a trial, where noninferiority is demonstrated in a study designed to show is not recommended by European guidelines for noninferiority studies The of the control arm There needs to be data to show that the arm is in an with over or current standard of care. In addition, the of the arm needs to be similar to that seen in trials This is to the chance of the that may be seen if noninferiority trials are that each use as a the regimen shown to be noninferior in the trial. For treatment of naive patients, a recent has improvements in the efficacy of HAART over the past 10 1. is important to new treatments with control arms that have shown the strongest efficacy, and have not been For example, once lopinavir/ritonavir showed efficacy over in the Abbott trial was no longer used as a control arm. In treatment-experienced patients, new treatments with efficacy are to the background which should to incremental improvements of efficacy in control arms. The efficacy of control arms has improved between the and trials, as the use of and then darunavir/ritonavir was in the control arms There are of noninferiority trials with control arms which are either not approved by or no longer recommended in treatment For example, the trial was powered to show noninferiority of fosamprenavir/ritonavir versus a control arm of which is no longer recommended for first-line When trials are it is difficult to the future standard of at the time the results be In situations this, it may be to conduct trials, to drugs new of care. The KLEAN trial the efficacy of fosamprenavir/ritonavir a more control, lopinavir/ritonavir, after the results of the trial were The TITAN trial was a of darunavir/ritonavir mg twice daily versus lopinavir/ritonavir in treatment-experienced patients 12. Figure 2 shows a review of trials to show that the control arm of the TITAN trial, lopinavir/ritonavir, in a similar to trials in patients. This approach can be used to show the of the efficacy in the control arm. showed higher efficacy than for naive patients in the Abbott trial and higher efficacy than control for patients in the Abbott trial The efficacy of the lopinavir/ritonavir arm of the TITAN trial similar to the efficacy of lopinavir/ritonavir in other trials of patients which the noninferiority for darunavir/ritonavir in the TITAN 2: HIV RNA 50 copies/ml at week 48 in the TITAN trial compared with lopinavir viral time to loss of virological trials in treatment-experienced patients The standard design for trials in highly patients has been to use an arm for all patients and then to to use or not use a new experimental These trials have been designed to show an efficacy benefit for the new This was used for the trials of the trials of maraviroc the trials of and the trials of The and trials of and tipranavir included in the control arms, but resistance predicted efficacy for the control chosen These trial have been given the excess of virological failure in the control arm, and the of drug which could future treatment of new with a low for are likely to to suppression of HIV RNA in the of treatment-experienced patients. For example, data from the trials of the showed at least of patients with HIV RNA 400 copies/ml when was either with or both drugs 5. Given these new developments, there may be in new clinical trials in which a control arm is expected to in the most treatment-experienced patients. In future noninferiority trials could be used to drug combinations which a high efficacy rate in treatment-experienced patients, but with a lower pill lower of drugs required and adverse For example, patients with current suppression on more combinations as could be randomized either to their current or to a more combination of new drugs. patients could be combinations of new and randomized to either or of their background regimen as These noninferiority trial are less likely to to excess virological failures in control arms, and could to new treatments for the of trial the efficacy of drugs is compared in patients also given an background of other it is important to the to which the background used would the efficacy and to the randomized treatment is to the efficacy Summary and for trial analysis and interpretation studies are designed to show similar efficacy in the experimental and control arms. Therefore, this of design should be adopted for new trials of treatment-experienced patients, in to trials powered to show differences versus the control arm. trials have the disadvantage of the control to a higher of virological failure, which may no longer be given the of new treatment The HIV RNA 50 endpoint should be used as the primary endpoint The primary and endpoints of a noninferiority trial to be in the protocol, with a for the control arm used, and statistical procedures to use in unexpected superiority is seen for the experimental versus the control arm. A standardized ‘per protocol’ analysis needs to be with similar exclusions across clinical As a the following patients should be excluded from the per protocol patients not take at least one dose of randomized patients the randomized patients taking an background or take experimental drugs the trial a defined The results from and analysis should be consistent to noninferiority. the intent-to-treat and per protocol analyses of noninferiority trials may be dominated by nonvirological endpoints. A ‘nonvirological failures censored’ should be including all virological failures while patients are receiving randomized and data after treatment discontinuation for nonvirological These ‘nonvirological failures censored’ analyses could treatments which are inferior but than the control arm. of new randomized trials designed to show noninferiority to be to show whether this has or has not been shown. The design and delta should be in the Where the 95% confidence interval for the true treatment overlaps but noninferiority was not the study results should not be interpreted as ‘no or Furthermore, the should be in the of the population including resistance noninferiority in naive patients may not the same result would in patients, and
Hill et al. (2008) studied this question.