Coronary heart disease (CHD) is a prevalent and life-threatening chronic cardiovascular condition, and early screening and diagnosis are essential for improving patient outcomes. Traditional diagnostic methods, such as electrocardiography, echocardiography, and coronary angiography, typically require specialized equipment, expert interpretation, and in some cases, invasive procedures, thereby limiting their accessibility and scalability for population-wide screening. In this work, we propose a novel, painless, non-invasive, and rapid auxiliary diagnostic approach for CHD based on scleral images. Specifically, we present the ScleraMIL model, a multi-instance learning framework that integrates the strengths of convolutional neural networks (CNNs) and Vision Transformer (ViT) architectures to capture both local and global representations across multiple scleral images from each individual. To ensure the model focuses on scleral features, the Mamba-UNet segmentation model is first employed to precisely extract the scleral region from the raw ocular image. Then each subject’s ten segmented scleral images (five per eye) are treated as instances forming a bag with a single patient-level label. These instances are processed through a CNN-based feature extractor, followed by a Transformer encoder that models inter-image dependencies for discriminative feature fusion, and finally passed to a classification head to predict the probability of CHD. Experimental results demonstrate that ScleraMIL achieves superior performance on key metrics such as accuracy, AUC, F1-score, and recall, significantly outperforming other deep learning and multiple instance learning methods. This work explores the feasibility of using scleral imaging as a potential biomarker for CHD detection and represents a promising step toward intelligent, accessible, and non-invasive cardiovascular risk assessment.
Gao et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: