Background/Objectives: Feature selection approaches should satisfy all evaluation criteria required by state-of-the-art chemoinformatic models. Our aim is to develop a methodology that is robust, interpretable and computationally efficient. Methods: This study presents a robust methodology for developing highly interpretable and computationally efficient predictive models, with a specific application in the discovery of SARS-CoV-2 main protease inhibitors. We evaluated various descriptor selection procedures to identify a transparent and reproducible approach that provides actionable insights for data-driven decisions. The models were trained and tested using molecules from the CHEMBL database and further validated on an external set of compounds. Results: Our findings demonstrate that a recently proposed procedure, combining the FeatureWiz algorithm with stepwise feature selection, is the only approach that satisfies all evaluation criteria required by state-of-the-art chemoinformatic models. In particular, we found that models based on two-dimensional descriptors and Ordinary Least Squares regression achieved the best results. Conclusions: Our framework and the choices made offer significant advantages in a decision-making context due to their inherent interpretability and computational efficiency. Our derived models, benchmarked against those in the literature, serve as effective, transparent tools for the rapid and reliable prediction of biological activity, providing a validated framework for data-driven decisions in drug discovery and beyond.
Stanković et al. (Wed,) studied this question.