Key result
PatchECG achieves state-of-the-art arrhythmia detection with ~5 times greater computational efficiency than 2D models.
Why the study?
Cardiovascular diseases are the leading cause of death worldwide, and accessible ECG monitoring creates a need to leverage large-scale unlabelled data via self-supervised learning to improve automated arrhythmia detection while reducing overfitting to class imbalance and noise.
Does PatchECG, a 1D Transformer model with self-supervised learning, improve automated cardiovascular arrhythmia detection compared to existing models?
Does PatchECG, a 1D Transformer model with self-supervised learning, improve automated cardiovascular arrhythmia detection compared to existing models?
Effect estimate: AUROC 0.9660 (95% CI 0.963-0.969)
Absolute Event Rate: 0.966% vs 0.9635%
A novel 1D Vision Transformer model (PatchECG) pre-trained on 8.2 million ECGs using self-supervised learning achieves state-of-the-art arrhythmia detection performance while being highly computationally efficient.
Advances reliable automated arrhythmia detection in clinical ECG workflows; extends self-supervised pre-training of Vision Transformers on large unlabelled datasets.
Cardiovascular diseases are the leading cause of death worldwide. With electrocardiogram (ECG) machines becoming more accessible, passive monitoring for arrhythmia detection is now possible. This work highlights the importance of self-supervised learning in detecting arrhythmias by leveraging large-scale unlabelled ECG data to improve performance and reduce overfitting to class imbalance and noise. We propose Masked Patch Modelling (MPM) and use 8.2 million unlabelled ECGs for self-supervised pre-training, introducing PatchECG, a 1D Transformer model that can be fine-tuned for various ECG tasks. PatchECG achieves state-of-the-art results on standard datasets, including PTB-XL multi-label classification, and sets new benchmarks on the largest and highest-quality multi-label dataset to date. Compared to existing methods, PatchECG is five times more computationally efficient while increasing model capacity by a factor of 14. We also compare the 1D PatchECG model to a state-of-the-art 2D vision Transformer, HeartBEiT, and observe significantly higher performance. Finally, ablation studies reveal a 2% performance improvement in handling class imbalance, label noise, and over-parameterization. These findings demonstrate the potential of self-supervised learning in advancing automated arrhythmia detection.
No takes yet. Share an insight, caveat, or question.
Chatterjee et al. (2026) studied Cardiovascular arrhythmia (n=8,468,000). PatchECG (1-dimensional Vision Transformer with Masked Patch Modelling) vs. 2D Vision Transformers (HeartBEiT) and CPC-based LSTMs was evaluated on Macro-AUROC for multi-label arrhythmia classification on the Unified Dataset (AUROC 0.9660, 95% CI 0.963-0.969). PatchECG, a 1-dimensional vision transformer pre-trained on 8.2 million unlabelled ECGs, achieved state-of-the-art arrhythmia detection (AUROC 0.9660) while being 5 times more computationally efficient than 2D models.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: