PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 17, 2025SHILAP Revista de lepidopterología16 citationsOpen Access

Leveraging high-throughput molecular simulations and machine learning for the design of chemical mixtures

ACAlex K. ChewMAMohammad Atif Faiz AfzalZKZachary Kaplan

Key Points

  • To develop and validate computational machine learning workflows for navigating the vast chemical design space and accurately predicting the physical properties of solvent mixtures.
  • Generated a dataset of over 30,000 solvent mixtures using high-throughput classical molecular dynamics simulations.
  • Evaluated three machine learning architectures connecting molecular structure and composition to properties: formulation descriptor aggregation (FDA), formulation graph (FG), and a Set2Set-based method (FDS2S).
  • Assessed model transferability against real-world experimental datasets from energy, pharmaceutical, and petroleum applications.
  • The Set2Set-based approach (FDS2S) outperformed both formulation descriptor aggregation and formulation graph models in predicting simulation-derived properties.
  • The predictive models identified promising chemical formulations at least 2 to 3 times faster than random guessing.
  • Trained models demonstrated robust transferability by accurately predicting experimental mixture properties across energy, pharmaceutical, and petroleum domains.

Abstract

Mixtures of chemical ingredients, such as formulations, are ubiquitous in materials science, but optimizing their properties remains challenging due to the vast design space. Computational approaches offer a promising solution to traverse this space while minimizing trial-and-error experimentation. Using high-throughput classical molecular dynamics simulations, we generated a comprehensive dataset of over 30,000 solvent mixtures to evaluate three machine learning approaches that connect molecular structure and composition to property: formulation descriptor aggregation (FDA), formulation graph (FG), and Set2Set-based method (FDS2S). Our results demonstrate that our new FDS2S approach outperforms other approaches in predicting simulation-derived properties. Formulation-property relationships can reveal important substructures and identify promising formulations at least two to three times faster than random guessing. The models show robust transferability to experimental datasets, accurately predicting properties across energy, pharmaceutical, and petroleum applications. Our research demonstrates the utility of high-throughput simulations and machine learning tools to design formulations with promising properties.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chew et al. (2025) studied this question.

synapsesocial.com/papers/69debe931d9bba5129b0cba0https://doi.org/10.1038/s41524-025-01552-2
Ask AI
Helpful
Bookmark
Share
View Full Paper