Key result
An SVM classification model using 72 selected descriptors and ECFP_4 structural fingerprints achieved an accuracy of 0.984 and kappa of 0.733, outperforming commercial hERG prediction models.
Why the study?
Most machine learning models predicting hERG inhibition rely on datasets from a single database, limiting their predictability and scope.
Absolute Event Rate: 0.984% vs 0.905%
A machine learning model trained on a massive integrated database of over 291,000 compounds demonstrated high accuracy in predicting hERG liability, a key preclinical safety screening step for proarrhythmic risk.
Supports refined hERG screening for proarrhythmic risk in drug development; extends computational models but leaves open prospective validation.
Assessing the hERG liability in the early stages of drug discovery programs is important. The recent increase of hERG-related information in public databases enabled various successful applications of machine learning techniques to predict hERG inhibition. However, most of these researches constructed the datasets from only one database, limiting the predictability and scope of the models. In this study, a hERG classification model was constructed using the largest dataset for hERG inhibition built by integrating multiple databases. The integrated dataset consisted of more than 291,000 structurally diverse compounds derived from ChEMBL, GOSTAR, PubChem, and hERGCentral. The prediction model was built by support vector machine (SVM) with descriptor selection based on Non-dominated Sorting Genetic Algorithm-II (NSGA-II) to optimize the descriptor set for maximum prediction performance with the minimal number of descriptors. The SVM classification model using 72 selected descriptors and ECFP_4 structural fingerprints recorded kappa statistics of 0.733 and accuracy of 0.984 for the test set, substantially outperforming the prediction performance of the current commercial applications for hERG prediction. Finally, the applicability domain of the prediction model was assessed based on the molecular similarity between the training set and test set compounds.
No takes yet. Share an insight, caveat, or question.
Ogura et al. (2019) studied hERG inhibition (in silico prediction) (n=291,219). Support Vector Machine (SVM) model with 72 selected descriptors and ECFP_4 vs. Commercial prediction models (ACD/Percepta, ADMET Predictor, StarDrop) was evaluated on Prediction accuracy for hERG inhibition on the test set. An SVM classification model using 72 selected descriptors and ECFP_4 structural fingerprints achieved an accuracy of 0.984 and kappa of 0.733, outperforming commercial hERG prediction models.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: