AI-ECG HCM model outputs in 681 patients correlated significantly with maximal wall thickness (r = 0.30), NT-proBNP (r = 0.41), and T-wave axis (r = 0.48), preserving biologically relevant signals.
Responsible external validation of AI-ECG tools requires appropriate study design, including control groups and validated operating thresholds, to accurately assess model performance and avoid misinterpretation.
We read with concern the study by Babur Guler and colleagues, which evaluated three AI-enhanced electrocardiography (AI-ECG) tools in 681 patients with confirmed hypertrophic cardiomyopathy (HCM).1 These tools include one for detecting underrecognized HCM, PRESENT-SHD, for screening for structural heart disease (SHD), and a multilabel rhythm and conduction classifier (ECGDx). These models developed by our group are freely available for research use on our lab’s website.2–5 We are committed to open and transparent research and welcome independent external evaluations. We agree with the authors that independent assessment of AI tools in disease-specific populations can inform understanding of the model’s behaviour in specific clinical scenarios, which is essential for the field. However, responsible external validation carries a reciprocal responsibility: Models must be represented accurately with respect to their intended use case, validated operating thresholds, and prior evidence base, and evaluated using an appropriate study design. On several of these fronts, the current study falls short in ways that undermine its conclusions. First, and most fundamentally, the study includes no control or reference group. All three tools are discriminative classifiers, developed and validated in mixed case-control populations. Applying them to a cohort composed entirely of confirmed HCM patients makes it mathematically impossible to assess diagnostic performance. The authors observe that ‘when applied to an exclusively HCM population, the tool appears to lose discriminative power, possibly due to the absence of the contrast provided by non-HCM cases in the training data’. This interpretation is fundamentally flawed. The apparent loss of discriminative power is not an empirical finding but a mathematical inevitability of applying a discriminative classifier to a case-only cohort. While the paper repeatedly frames these distributions as ‘modest performance’, sensitivity, specificity, and discrimination cannot be estimated from a case-only sample, and the resulting probability distributions cannot be benchmarked against any meaningful clinical reference. Second, the study does not use the validated operating thresholds from the primary publications. The validated thresholds are 15% for the HCM model and 20% for PRESENT-SHD, each selected to achieve approximately 90% sensitivity in internal validation.3–5 Instead, it emphasizes proportions of patients with HCM or SHD probabilities above 50% and 75%, which represent arbitrary cut-offs.1 As a result, the conclusion that relatively few patients ‘receive a high score’ is driven by threshold choice, which conflates a study design decision with model failure. Third, applying screening tools developed in treatment-naive populations to a mixed post-intervention cohort without stratification introduces avoidable confounding. This cohort includes patients who had undergone ventricular myectomy (2.5%), alcohol septal ablation (5.5%), and implantable cardioverter-defibrillator implantation (11%), representing interventions known from previous work to substantially alter the electrocardiographic HCM signature. Our multicentre evaluation demonstrated that AI-ECG HCM scores do not decrease, and may paradoxically rise after myectomy, reflecting potential scarring from the procedure.5 In clinical practice, there is limited utility for an HCM screening tool being used in a post-intervention population with an established diagnosis. Although evaluating model outputs in such individuals may still be informative for understanding model behaviour or exploring whether these scores capture features of disease severity, such analyses should not be used to draw conclusions about discriminative performance without appropriate stratification and interpretation. Fourth, deployment details are insufficiently documented. No ECG images are shown to confirm that pre-processing, including cropping and de-identification, aligned with validated input specifications. Finally, no information is provided on when ECGs were acquired relative to diagnosis or treatment, making pooled results difficult to interpret across what may be vastly different stages of the disease trajectory. Importantly, the results cited as evidence of underperformance could equally be interpreted as evidence that these models retain meaningful biological signal in a highly complex, disease-enriched cohort. Figure 1 of the manuscript shows that, at prespecified operating thresholds, most individuals in this cohort would in fact have been classified as having HCM. Moreover, even in a population outside the model’s intended screening setting, HCM model outputs correlated significantly with maximal wall thickness (r = 0.30), NT-proBNP (r = 0.41), late gadolinium enhancement, T-wave axis (r = 0.48), and Sokolow-Lyon index (r = 0.39). Apical HCM, a phenotype characterized by marked repolarization abnormalities, yielded the highest model probabilities. Likewise, PRESENT-SHD probabilities correlated inversely with LVEF measured by echocardiography (r = −0.22) and CMR (r = −0.16), as well as with TAPSE (r = −0.24), aligning with the functional parameters a structural heart disease screening model would be expected to capture in a population with established cardiomyopathy. Thus, the models continued to track clinically meaningful gradients of disease expression, despite this setting not being designed to assess diagnostic discrimination. These findings support preservation of biologically relevant signal rather than model failure. The misinterpretation of publicly available AI-ECG tools, as illustrated in the present study, has implications that extend well beyond any individual model. Open science in cardiovascular AI remains uncommon, and its continued progress depends on a culture of collaborative, methodologically rigorous validation.6–8 When openly shared research tools are applied outside their intended use cases, evaluated without appropriate comparators, and then framed as underperforming despite preserving meaningful biological signal, the consequences extend beyond local misinterpretation: such practices risk discouraging future model sharing. Yet this problem is readily preventable. Just as commercial AI diagnostics typically come with implementation guidance and technical consultation to promote appropriate deployment, openly shared research tools can also be evaluated responsibly.9 In many cases, brief communication with corresponding authors would be sufficient to clarify intended use cases, validated thresholds, and key study design considerations. We raise these points not to restrict access or diminish the investigators’ considerable effort, but to emphasize that this exceptional cohort with evident scientific depth deserved a study design equal to its quality. To support future evaluations and provide a more generalizable framework for the field, we propose a checklist for responsible external validation of AI-ECG tools (Table 1). We welcome independent validation of our tools and would be glad to collaborate on a rigorous analysis of this registry, because fair and methodologically sound evaluation is essential not only for trust in individual models but also for preserving the culture of openness on which progress in cardiovascular AI depends.10 Checklist for responsible external validation of artificial intelligence enhanced electrocardiography tools Lovedeep Singh Dhingra (Conceptualization, Methodology, Writing—original draft lead), Philip M. Croon (Conceptualization, Methodology supporting, Writing—original draft equal), Evangelos K. Oikonomou (Conceptualization, Methodology supporting, Writing—review the collection, management, analysis, and interpretation of the data; the preparation, review, or approval of the manuscript; or the decision to submit the manuscript for publication.
Dhingra et al. (Thu,) conducted a letter in Hypertrophic cardiomyopathy (HCM) (n=681). AI-enhanced electrocardiography (AI-ECG) tools vs. No control or reference group was evaluated. AI-ECG HCM model outputs in 681 patients correlated significantly with maximal wall thickness (r = 0.30), NT-proBNP (r = 0.41), and T-wave axis (r = 0.48), preserving biologically relevant signals.