Background: Asian Americans (AAs) experience disproportionately higher rates of intracerebral hemorrhage (ICH) compared to other populations. Accurate stroke subtype prediction remains challenging due to complex risk factor interactions that traditional statistical approaches may miss. Early subtype identification optimizes treatment protocols and secondary prevention. This study evaluated machine learning (ML) feasibility for stroke subtype classification among AA patients. Methods: We analyzed electronic health records (EHR) from TriNetX with 54,948 AA stroke patients (2010-2019) across 55 US healthcare organizations. Patients were classified into five stroke subtypes: ischemic stroke (IS), ICH, subarachnoid hemorrhage (SAH), transient ischemic attack (TIA), and multiple stroke types. Predictors included demographics, vitals, lab values, relevant comorbidities and medications. We applied random forest (RF) imputation for missing values, inverse-frequency weights for class imbalance, and Classification and Regression Trees (CART) and RF algorithms with 10-fold cross-validation on 80/20 train-test splits. Results: The cohort comprised 67.2% IS, 11.2% ICH, 10.4% SAH, 6.6% TIA, and 4.6% patients with MS. Mean age was 61.6±14.1 years with 53.4% male. Hypertension (76.9%), hyperlipidemia (61.7%), and diabetes (39.1%) were most prevalent comorbidities. RF outperformed CART with 72.8% vs 68.4% accuracy and 0.875 vs 0.774 AUC. Variable importance analysis identified stroke history (SAH, ICH, IS) as the strongest predictors, followed by diastolic blood pressure variability, HDL cholesterol, and mean diastolic BP. Both models showed high specificity but moderate sensitivity. Conclusion: ML algorithms classified stroke subtypes in AAs with good accuracy and clinical feasibility for improving acute care delivery. Stroke classification directly impacts treatment selection, as different subtypes require distinct therapeutic approaches and have varying prognoses. Our finding that stroke history emerged as the strongest predictor highlights prior cerebrovascular events' importance in subtype determination. Yet, our current EHR-based stroke history definitions (diagnosis before the index event by ≥1 day) may misclassify concurrent events as historical and affect interpretability. ML's ability to capture complex interactions makes it valuable where traditional approaches are inadequate. Future studies may consider temporal validation to develop robust clinical decision tools.
Ding et al. (Thu,) studied this question.