To the Editor: In the article by Glassman et al “Benefit of Transforaminal Lumbar Interbody Fusion vs Posterolateral Spinal Fusion in Lumbar Spine Disorders: A Propensity-Matched Analysis From the National Neurosurgical Quality and Outcomes Database Registry,”1 the authors attempt to use the retrospective observational, multicenter data from the National Neurosurgery Quality and Outcomes Database (N2QOD) to compare the outcomes of transforaminal lumbar interbody fusion (TLIF) vs posterolateral spinal fusion (PSF). Their analysis showed that, for spondylolisthesis patients, TLIF had greater functional outcomes up to 12 months. In the Methods section, the authors noted that “Propensity matching allows matching of multiple patient characteristics across groups without performing a one-on-one matching of each case to a control.”1 However, this is a statistical technique that has wide latitude for applicability2,3 and is, itself, controversial.4 Before we accept this technique as a sound method to parse N2QOD, I think we should understand it better. A propensity score is defined as the probability of assignment to treatment vs control (in this case, PSF vs TLIF) that is calculated from a set of observed variables, which are thought to be confounders, otherwise known as covariates.2,3,5 This has become somewhat popular as a method of adjusting for known confounding factors in observational studies,6 in effect creating a “quasi-randomized” experiment.3 Here, we have a small group (NPSF = 306) and a large one (NTLIF = 1230). First, the authors decided to consider the TLIF group as the “control” group and the smaller, PSF group as the “treatment.” Then they subclassified the cohort into 3 subgroups: spondylolisthesis, adjacent segment disease, and spinal stenosis; to maintain concision, I will restrict discussion to the spondylolisthesis group. They identified 134 spondylolisthesis patients who had PSF, and 637 who had TLIF. The variables I have shown in Table 1 were listed in the Methods section by the authors as the known covariates.1 Matching is, in essence, the search for a dataset that might have resulted from an RCT but is hidden in a larger, observational data set.4 In a case–control study, for example, each of these variables would ideally be matched to reduce bias.7 The trouble arises (and is magnified by N2QOD) when the number of variables becomes very large. In a “homegrown study,” I may be able to match a case to a control for age, body mass index, and levels fused. But tack on all these other variables in Table 1 and the task becomes hopeless even with a large cohort. To solve this problem, we make a compromise. We assume that the variables in Table 1, the known confounders, are all that is needed to make the decision on whether to perform a PSF as opposed to a TLIF. Since each variable may have a different effect on the decision to treat, we construct a single scalar variable,3 the propensity score, from a mathematical function that depends on all these individual variables, for each patient. How? The propensity score is, in effect, the calculated probability the patient will get a PSF given only the variables in Table 1. We use the actual treatment status in a logistic regression analysis with the covariates as the independent variables to create a regression model that will give us a propensity score for any patient with the collected covariates.3,5 Table 1. - Variables used in Propensity Matching (Covariates) Age Sex Body mass index Smoking status Race Educational level Employment status Anxiety disorder Insurance status Workers’ compensation status Symptom duration ASA grade Number of levels fused Preop back pain NRS Preop leg pain NRS Oswestry disability index EuroQOL-5D EuroQOL-5D, EuroQol-5-Dimension; NRS, Numeric pain Rating Scale The matching algorithm is, essentially, as follows.3,5 First, we calculate the propensity score for each of the 134 PSF and 637 TLIF patients based on their Table 1 variables that we have collected from N2QOD. Second, we randomize the 2 lists of patients. We pick the first PSF patient and work down the TLIF list until we find a patient with the same propensity score. We then imagine that these 2 patients, because they had the same propensity score, were equally likely to have either the treatment (PSF) or the control (TLIF). They were thus “randomly” assigned: one to treatment (PSF) and the other to control (TLIF). We take both patients out of the list and iteratively repeat until each PSF patient has a corresponding TLIF patient and out dataset is complete. Very Clever…Now, the problems. First problem: only 109 of the 134 PSF patients had a match. That leaves 25 patients (18% of the group) unaccounted for. What happened? The authors didn’t say but if their propensity scores didn’t match to any of the 528 remaining patients in the TLIF group it hints at covariates that were not in the model. Second problem: now we have 109 treatment and 109 control patients with, thanks to propensity matching, in the authors’ words, “no significant differences in baseline demographics or health status measures between the PSF and TLIF groups.”1 Well, when the statistics software used those same variables to create these 2 groups, I’m sure it was able to find that the 2 groups were alike. But it didn’t! The data summary for the entire cohort, shown in their Table 2,1 showed significant differences in smoking, American Society of Anesthesiologist (ASA) grade, and number of levels fused. There was also a significant difference in estimated blood loss but that was not a propensity-matched variable. Now, the statistics software used a propensity matching algorithm to pluck out 109 patients for each group out of this cohort and there was still a significant difference in number of levels fused (Table 3, P = .021, Table 4, P = .034)1 between the PSF and TLIF group. What this means, essentially, is that the propensity matching algorithm was able to eliminate the differences between the PSF and the TLIF group when it came to factors like ASA grade and smoking, but the groups were too different when it came to number of levels fused to say that this “quasi-randomization” really worked. Third problem: most of these variables aren’t really relevant. Again, propensity matching is a technique used to reduce bias by identifying confounding factors and eliminating the possible bias that can be attributed to them. It is considered weaker than a randomized controlled trial (RCT), for example, because, for it to work, you have to “know” all the confounding variables whereas an RCT can theoretically account for “unknown” ones as well. But how do variables like sex, race, or ASA grade help decide whether to perform an TLIF vs PSF? The authors included these for only 1 reason: because they could. Lost in this flood of unlikely confounders, we are ignoring the variables that probably are contributing to the differences between the TLIF and the PSF groups. Even the authors tacitly concede this. Their introduction mentions that TLIF theoretically offers “more complete foraminal decompression, better correction of deformity, and more effective treatment of discogenic pain.”1 So shouldn’t they add variables to account for these? They didn’t because they couldn’t—N2QOD doesn’t collect these. Fourth problem: because the comparability between the groups is bad, they drew bad conclusions when common sense should have told them to do otherwise. The authors noted that they were “surprised” to find similar operating room times and estimated blood loss between TLIF and PSF. The discussion then jumps to the conclusion that this is the result of “improvement in TLIF technology” and “progressive adoption of minimally invasive surgical techniques” but they conveniently forgot the most likely explanation, supported by the study, which was that the PSF group had a statistically greater percentage of multilevel fusions. More levels = more time and more blood loss. What is the big picture here? Let's postulate that there are 3 types of surgeons: surgeon X only does PSF, surgeon Y does only TLIF, and surgeon Z does both. There are 3 clinical questions that I would like to know from a comparison of TLIF vs PSF. Question 1: How should surgeon Z choose between TLIF vs PSF on any given patient? Only Z cares about this as he necessarily has some criteria to choose one method over another on any given patient. Question 2: For any given patient, which is better, TLIF or PSF? That is, if we put surgeon X up against surgeon Y with the same types of patient, who does better? Question 3: Does surgeon Z, using his specified criteria, do better than either X or Y? This study answers none of these questions yet it pretends to. By using propensity matching, the authors are essentially saying that, for question 1, surgeon Z just needs to look at the variables in Table 1 because they are all that mattered in their assignment of patients to PSF vs TLIF group. I know of no surgeon Z who will ignore issues of radiological structural findings, deformity, bone quality, or previous surgery but takes into consideration things like anxiety disorder or educational level. What about questions 2 and 3? Well, good RCTs could answer them but they have been rare.8 This study1 is an attempt to approximate an RCT by getting every variable they could collect and using them to match patients to give us a false impression that bias has been minimized to the best of their ability. But, as I hope to have demonstrated, these groups are not comparable. The reason they are not comparable, if I may hazard a guess, is because they are not comparable. That is, these patients essentially have subtly different pathologies based on factors such as deformity, foraminal stenosis, and disc degeneration that are being skirted over by the Procrustean dataset that is N2QOD. That the authors used a controversial statistical technique2,4 to try to create comparable groups and it still didn’t work is testament to the severe limitations of using a retrospective, multicentered, commercial database to answer clinical questions that it was not designed for. If this dataset were at a single institution, it may have been possible to go back to the patient charts, revisit the data, hypothesize new variables, redesign the study, and try again. But here we are at a dead end. More patients in the N2QOD dataset are not going to fix it but will only make this problem worse. Disclosure The author has no personal, financial, or institutional interest in any of the drugs, materials, or devices described in this article.
No takes yet. Share an insight, caveat, or question.
Sandeep S. Bhangoo (2017) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: