Key points are not available for this paper at this time.
Molecular property prediction is at the core of machine learning (ML)-driven materials and drug discovery. Effectively navigating the ML workflow requires careful consideration of molecular representations, input preparation strategies, and model architectures. This review aims to provide an intuitive understanding of information processing and input construction in different machine learning models, offering guidance on how different ML approaches align with various prediction tasks and molecular representations. The methodologies are broadly categorized into descriptor-based, string-based, and graph-based models, with greater emphasis placed on graph-based approaches due to their superior ability to capture molecular structure and spatial geometry. Within the string-based category, focus is given to transformers and large language models (LLMs), which are gaining increasing attention owing to their success in natural language processing (NLP). Among graph-based models, the emerging class of geometric graph neural networks (Geometric GNNs) is discussed in detail, as these models represent the current state of the art across multiple benchmark datasets in molecular property prediction.
Thameem et al. (Sat,) studied this question.