501 Background: Administrative claims provide valuable data on real-world oncology costs and patient experiences. However, the diagnosis codes found in claims often lack the detail necessary to study variations in outcomes and costs by cancer stage. Applying machine learning methods to infer cancer stage from claims data could broaden the scope of population-level research. This study aimed to develop and validate a machine learning algorithm capable of predicting patients’ liver cancer stage at diagnosis using only claims data. Methods: Incident liver cancer cases diagnosed between 2016-2017 were identified using the SEER-Medicare data. Patients with <1 month of Medicare Parts A/B/D enrollment in 2016-2017, <12 months of A/B/D enrollment prior to diagnosis, cancer-related treatment within one year of index or prior cancer diagnoses were excluded. Claims were flagged for evidence, frequency, and timing of cancer-related surgeries, anti-cancer therapies, radiation therapy, hospice, and death. These flags, along with demographics, frailty-related diagnoses, and nursing home residence, were tested as predictors of liver cancer stage (American Joint Committee on Cancer – AJCC, local (L), regional (R), and distant (D) – LRD). Analysis was conducted in R Statistical Software (v4.1.2; R Core Team 2021) using predictive multinomial logistic regression (nnet package; Venables and Ripley 2002). Models were trained on 70% of the sample and tested on the remaining 30%. Results: The initial model achieved 47.3% accuracy 95% CI: 43.2%-51.4% in predicting SEER-derived AJCC liver cancer stages 1, 2, 3A, 3B, 3C, 4A, and 4B (n = 586). Collapsing the stages into two groups (1/2 vs. 3/4) increased accuracy to 75.4% 95% CI: 71.7%-78.8%. Using LRD staging, the model accuracy reached 66.6% 95% CI: 62.6%-70.4%, peaking at 70.6% 95% CI: 66.8% - 74.3% when collapsed to a binary model (L vs. R/D). Based on prior work with other cancer types, models with accuracies of at least 69% can produce stage-cohorts with statistically comparable cost and utilization patterns to cohorts staged using tumor registry data alone. Conclusions: Machine learning algorithms offer a practical approach for determining liver cancer stage at diagnosis using claims data, facilitating more extensive population-level research into costs and clinical outcomes by stage. Model accuracy is dependent on staging type predicted and the alignment of said staging system with clinical practice patterns.
Smith et al. (Sat,) studied this question.