Key points are not available for this paper at this time.
The process of drug discovery is one of the most expensive, time-consuming, and high-risk endeavors in modern science. Translating initial scientific insights into safe and effective therapies, supported by genomics, structural biology, and computational chemistry, typically requires more than a decade and substantial financial investment. Machine learning (ML) has emerged as a powerful tool for improving efficiency across the drug discovery pipeline. By enabling the analysis of large and complex datasets, ML supports target identification, lead discovery, optimization, and prediction of preclinical and clinical outcomes. Its integration with experimental validation and automation is illustrated by recent advances such as protein structure prediction, AI-driven antifibrotic compound discovery, and antibiotic identification. Despite these advances, significant challenges remain. Model generalizability is limited by data scarcity, heterogeneity, and hidden biases. In addition, the translation of in silico predictions into clinically validated outcomes remains a major bottleneck, and regulatory acceptance is constrained by limited model interpretability. Ethical considerations, including data privacy, equitable representation, and the potential misuse of generative models, further complicate adoption. This review examines the applications of ML across the drug discovery pipeline, with a focus on translational and regulatory considerations. It also discusses emerging directions, including hybrid physics–AI approaches, multimodal foundation models, federated learning, and explainable AI. The effective integration of ML will depend on rigorous validation, interdisciplinary collaboration, responsible data governance, and alignment with regulatory frameworks.
El-Tanani et al. (Fri,) studied this question.