In this issue of Paediatric and Perinatal Epidemiology, Simmons and colleagues 1 use insurance claims data to estimate the effect of initiating midwife- versus physician-led prenatal care on the risk of adverse delivery outcomes. They conclude that midwife-led prenatal care is associated with a decreased risk of caesarean delivery and maternal infection, but an increased risk of postpartum haemorrhage. We commend the authors for expanding upon previous evaluations of midwife-led prenatal care by including all pregnancies with observed outcomes—not just deliveries—and leveraging g-methods within the target trial emulation framework to estimate the effects of real-world interventions. By leveraging these methods, the authors clearly articulate their target causal estimand and the extent to which key components are emulated in their study. This study has the potential to inform guidelines for prenatal care delivery in the U.S., particularly since midwives represent an important component of the obstetrical workforce in nations with better maternal and infant health outcomes 2. In this commentary, we discuss the interpretation of the study findings within the context of its strengths and limitations. We focus specifically on (i) exposure misclassification, (ii) the proposed interventions and (iii) future directions. Simmons et al. define prenatal care initiation as the first claim during pregnancy with ≥ 1 diagnosis or procedure code indicative of ongoing pregnancy that was billed under a midwife, obstetrician, maternal-foetal medicine specialist or family practitioner. This definition may be subject to substantial exposure misclassification, threatening the internal validity of the findings. As the authors note, there are financial incentives to billing midwife-led care under physicians for higher reimbursement, which could lead to underestimation of the prevalence of midwife-led prenatal care. Indeed, the observed 1.4% prevalence of midwife-led care is substantially lower than the 10% observed elsewhere 3. We agree with the authors that such misclassification is likely independent and non-differential with respect to the study outcomes and would bias findings towards the null in expectation (under the additional assumptions of no misclassification of other variables and independence with respect to covariates) 4. However, quantifying the potential bias is necessary for clinical and policy decision-making, particularly since this bias could explain the study's null findings for obstetrical trauma and admission to the neonatal intensive care unit. Simple quantitative bias analysis to correct for exposure misclassification has been well-described for standard epidemiologic methods 5. Briefly, such an approach requires specifying sensitivity and specificity under a given distribution that dictates the reclassification of study participants within a 2 × 2 table of the exposure and outcome. We could conduct this analysis using the unadjusted risk for caesarean delivery in Table 2 of Simmons et al.: 21.8% for midwife- versus 33.0% for physician-led prenatal care, corresponding to a risk difference of −11.2%. If we assume a 14% sensitivity (derived by assuming that the true prevalence of midwife-led prenatal care was 10%; that is,1.4%/10.0%) and 99.9% specificity, we derive corrected risks of 20.9% and 34.3%, respectively, for a risk difference of −13.1%. Corresponding analyses on other primary outcomes would also result in corrected risk differences that were further from the null: from 0.9% to 1.1% for primary postpartum haemorrhage, 1.1% to 1.3% for secondary haemorrhage and −2.6% to −3.0% for maternal infection. These analyses would suggest that correcting for exposure misclassification would indeed move the estimates away from the null by a small amount. It is important to note, however, that these sensitivity analyses fail to replicate the interventions investigated by Simmons et al. The corrected risk differences we derived above correspond to the population average treatment effect, or the expected risk difference in the study outcome if all members in the baseline population did or did not initiate midwife-led prenatal care. In contrast, Simmons et al. aim to estimate the expected differences in risk should increasingly more study participants initiate midwife-led prenatal care. Other sensitivity analyses are necessary to target these estimands. As previously described, the g-formula is a generalisation of standardisation that, under specific assumptions 6, estimates the counterfactual risks for all members of a study population under investigator-specified treatment strategies 6, 7. Stepwise procedures for implementing g-computation as described by Simmons et al. are shown in figure 1 7. If we believe that the true prevalence of midwife-led prenatal care is 10%, it is tempting to consider the risk estimates under the intervention where 11.2% of pregnancies (corresponding to a 10% increase) initiate midwife-led care as closer to the true risk. Under this assumption, we could derive corrected risk differences by treating that analysis as the referent. However, such an approach would be invalid. The outcome models (from step 1) would have been constructed using a misclassified exposure, making the subsequent risks subject to bias. A valid sensitivity analysis would require correcting misclassified exposures before building the outcome and covariate models. Following existing literature for correcting misclassified outcomes using likelihood-based approaches 6, such an approach could involve adding a step to the standard g-computation procedure where the user first corrects the exposure misclassification based on prespecified sensitivity and specificity values (Figure 1). Corrected exposures would then be used in all subsequent steps. Formalisation of statistical methods for conducting such an analysis and subsequent simulation studies is necessary to compare the two approaches. The magnitude of the interventions proposed in this analysis—increasing midwife-led prenatal care from 1.4% to 11.2%, 21.1% and 50.1%—may limit the usefulness of these findings. Causal inference relies on the assumption of positivity, which states that for every combination of measured baseline covariates, there must be a non-zero probability of receiving each level of treatment 8. Given that only 1.4% of the population were observed to have initiated midwife-led care, it is likely that there were some near positivity violations. Further, parametric g-computation is more sensitive to violations of assumptions of no measurement error and no model misspecification than conventional methods 9. Such interventions may also be unrealistic. Unrealistic interventions can lead to problems in estimation because the models must extrapolate far beyond the observed data. These interventions also may not align with what can be reasonably achieved in the near-term: increasing the prevalence of midwife-led prenatal care to 50.1% of pregnancies, for example, would likely require a larger midwife workforce than currently exists in the U.S. 2 Policymakers, then, should not expect such large benefits until training programmes can meet these needs. Notably, underestimation of the prevalence of midwife-led prenatal care may make some of the interventions more realistic than they initially appear. If the rate of midwife-led prenatal care is truly around 10% 3, increases to 11.2% or 21.1% may be reasonable. Electronic health record (EHR) data may be exceptionally useful for overcoming some limitations of the current analysis because (1) provider type may be less subject to the billing constraints imposed by claims and (2) there may be more detailed encounter and covariate data. Strength (1) will clearly address exposure misclassification, as may strength (2). Specifically, some EHR databases contain structured encounter type variables (e.g., ‘Initial’ or ‘Routine’ prenatal care), which may help standardise prenatal care definitions across participants. The richer set of clinical variables typically available in EHR data would also allow more robust confounding adjustment, as well as consideration of interventions that incorporate grace periods and time-varying and dynamic prenatal care strategies. We caution that EHR data is also subject to limitations. EHR data, for example, are limited to only those encounters that occur within the included healthcare systems, while claims data include all reimbursed claims regardless of provider 10. We suggest, then, that EHR-based analyses would be complementary to those conducted by Simmons et al. Simmons et al. highlight an important context in which g-computation paired with principled target trial emulation can be used to estimate the effects of real-world prenatal care interventions. The analytic decisions made by Simmons and colleagues reflect thoughtful consideration of the limitations of U.S. insurance claims data and are an important extension of previous work that only included deliveries. With the increasing usage of g-computation, attention to sensitivity analyses that investigate the impact of exposure misclassification on study findings is warranted. Further, analyses that leverage g-computation must focus on realistic interventions that are reasonably supported by their data. Finally, claims-based analyses subject to such important limitations may benefit from replication in data sources with detailed encounter data. E.C.C. was invited to write the commentary. C.D.L. and E.C.C. wrote the article together. This work was supported by the Eunice Kennedy Shriver National Institute of Child Health and Human Development, R01HD113685, R01HD114736. C.D.L. has received payment from Target RWE, Amgen and Regeneron Inc. for unrelated work. E.C.C. has nothing to disclose. The authors declare no conflicts of interest. Data sharing not applicable to this article as no datasets were generated or analysed during the current study.
Latour et al. (Sun,) studied this question.