Computational small-molecule drug development aims to refine candidate small molecules toward specific functional outcomes, typically through iterative cycles of experimental assays and cheminformatics modeling. However, small experimental batch sizes and high resource demands challenge these processes. Machine learning-based structure-activity models can guide molecular generation, but their predictive accuracy often declines when newly proposed compounds deviate too far from the distribution of previously tested molecules. While diversity in a set of generated molecules is commonly a desirable metric, unconstrained molecular generation also risks producing irrelevant or impractical candidates. To address these challenges, we develop and assess a procedural framework that generates novel molecules via local perturbations and refines them via an actively trained preferential graph neural network and a probabilistic ranking model. We show that predicting expected differences in structure-activity relationships is an easier task than predicting absolute scores and that it can be used to inform a molecular generative process with learned models. We benchmark our approach against established conditionally generative algorithms over numerous evaluation metrics. Our method consistently achieved improved signal precision and maintained consistent performance across the population demonstrating its ability to optimize target functions across multiple domains. These results suggest a practical and generalizable pathway for integrating machine learning into early-stage small-molecule drug development.
Dreisler et al. (Sun,) studied this question.