DNA-encoded library (DEL) technology is a novel experimental drug discovery technique that merges combinatorial chemistry with next generation DNA sequencing to build small molecule libraries that are larger and easier to screen than traditional chemical libraries. Recently, DELs have been successfully utilized to discover novel small molecule drugs for important biological targets like sEH and WDR91. However, DELs inherently lack chemical diversity and drug-likeness, and the biological activity data obtained from DEL screens are redundant and noisy. Computational methods can be used to denoise data and extrapolate conclusions to wider, more diverse, and more drug-like sets of chemical space. Here, we present physics and machine learning based computational methods for analyzing and utilizing DELs for drug discovery. In this work, we build upon the Kaggle competition hosted by Leash Biotech and on the data set released with the competition, the big encoded library for chemical assessment (BELKA). Using the BELKA data set, we trained machine-learning models to predict the probability of small molecule binding to three different protein targets. We systematically determined the optimal machine-learning architectures, hyperparameters, and training splits to optimize prediction accuracy on out of distribution test sets. To supplement our machine-learning models, we used physics-based docking software to improve our predictive accuracy on hard-to-predict regions of chemical space. Finally, we used our models to predict binding of molecules from the Enamine REAL space to the three targets, and validated predictions using molecular dynamics. Our work provides insight into how to best integrate DELs and computational methods for efficient and novel drug discovery.
Marissa Dolorfino (Sun,) studied this question.