The multimodal ECG+echocardiogram model achieved AUROC 0.94 for LVSD, outperforming ECG-only (0.89) and echo-only (0.89) models in cardiovascular disease classification.
Does a multimodal foundation model integrating ECG and echocardiogram data improve diagnostic performance for cardiovascular conditions compared to unimodal models?
A multimodal deep learning foundation model integrating ECG and echocardiography data outperforms unimodal models in diagnosing cardiovascular conditions such as LVSD and atrial fibrillation.
Absolute Event Rate: 0% vs 0%
Abstract Background Clinicians leverage information from multiple data sources, such as the electrocardiogram (ECG) and echocardiogram to diagnose and manage cardiovascular disease. Deep learning models have shown immense potential to enhance these clinical processes by capturing digital biomarkers of cardiovascular disease from diagnostic modalities. However, the development of multimodal algorithms capable of making joint inferences from multiple modalities, mirroring the clinicians’ decision-making, has remained an unmet need. Purpose We aimed to develop a multimodal foundation model of ECG and echocardiogram as an integrative deep learning approach to evaluating cardiovascular conditions. Methods Twelve-lead ECGs and transthoracic echocardiograms performed at a large academic health system between 2001-2021 were collected. Individuals with paired ECGs and echocardiograms recorded within 3 years were identified and the data was split at the individual level into training, validation, and test sets (80/10/10%). Using paired ECGs and echocardiograms, we developed a deep learning model in a self-supervised setting with contrastive loss and merged embeddings in a joint latent space (Figure 1). We evaluated the performance of this foundation model after fine-tuning for left ventricular systolic dysfunction (LVSD defined as left ventricular ejection fraction of 40% or less), atrial fibrillation, age, and sex as auxiliary tasks. The performance of the multimodal model was compared with those of supervised unimodal ECG-only and echocardiography-only models for each task. Results In total, 4,391,610 videos from 129,424 echocardiography studies paired with 1,094,656 ECGs from 53,876 unique individuals were included (67 ± 14 years of age, 43% female, Figure 1). The multimodal algorithm demonstrated excellent discrimination for disease classification in the test set. In classification of LVSD, for example, it achieved an area under the receiver operating characteristic (AUROC) curve of 0.94, area under precision recall curve (AUPRC) of 0.99 among individuals with LV ejection fraction greater than 40%, and AUPRC of 0.72 among those with LV ejection fraction of 40% or less. The multimodal model consistently outperformed the unimodal algorithms, including the ECG-only and echocardiography-only models in classification of LVSD (AUROC ECG+Echo 0.94, ECG 0.89, Echo 0.89), atrial fibrillation (AUROC ECG+Echo 0.95, ECG, 0.94, Echo 0.75), sex (AUROC ECG+Echo 0.93, ECG 0.87, Echo 0.78) and age (R for variance ECG+Echo 0.75, ECG 0.67, Echo 0.58, Figure 2). Conclusion We have developed a multimodal foundation model for joint inference from ECGs and echocardiograms for diagnosis of cardiovascular conditions. This novel approach outperforms the conventional unimodal models in disease diagnosis and represents an advance in integrating the clinical applications of artificial intelligence.Figure 1 Figure 2
Nargesi et al. (Sat,) reported a other. The multimodal ECG+echocardiogram model achieved AUROC 0.94 for LVSD, outperforming ECG-only (0.89) and echo-only (0.89) models in cardiovascular disease classification.