Quantitative structure–activity relationship (QSAR) modeling underpins computational drug discovery, yet the factors governing model generalizability across structurally novel compounds, and across protein families, remain incompletely understood. This study addresses that gap through systematic 2D-QSAR benchmarking across four therapeutically diverse ChEMBL targets, EGFR (kinase, n = 10, 036), DRD2 (GPCR, n = 7, 558), BACE-1 (protease, n = 8, 080), and hERG (ion channel, n = 8, 077), totaling 33,751 compounds under identical experimental conditions. Three principal findings emerge. First, scaffold-based splitting revealed target-dependent generalization gaps varying 2-fold across targets (mean gap: 0.084–0.171 R 2 units across five random seeds), demonstrating that the interpolation-to-extrapolation performance penalty is a data set property, not an algorithmic limitation. Second, applicability domain (AD) analysis confirmed substantial prediction quality degradation outside the Tanimoto-based chemical domain (0.31–0.51 R 2 drop), with hERG showing near-total loss of predictive power for structurally distant compounds ( R 2: 0.62 → 0.11). Third, algorithm rankings were remarkably consistent across all four protein families (mean Spearman ρ = 0.92), with tree-based ensemble methods, particularly Random Forest, on ECFP4 fingerprints providing the most consistent performance (scaffold-split R 2: 0.40–0.59). A basic 3-layer graph convolutional network (without edge features or pretraining) underperformed RF + ECFP4 on every target (mean deficit: −0.22 R 2 ). These results establish reproducible prospective baselines and demonstrate that AD reporting is an essential, not optional, component of QSAR model evaluation.
Alshehade et al. (Tue,) studied this question.