PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 18, 2026Journal of Chemical Information and Modeling0 citations

Multiscale Hypergraph Masked Autoencoder with Δ-Property Alignment for Novel Molecular Representation Learning

View Full Paper
ZZZiyan ZhuYWYang WangXWXian Wei

Key Points

  • To improve molecular representation learning by capturing cross-scale interactions and functional group semantics using a hypergraph framework.
  • Proposed a Multiscale Hypergraph Convolutional Masked Autoencoder (MSHG-MAE) model for molecular representation.
  • Utilized atom nodes and multitype hyperedges to represent molecules as hypergraphs.
  • Employed a semantics-aware masked autoencoding objective for biased masking of atoms and hyperedges.
  • Implemented Δ-Property Alignment to link embedding differences with proxy property differences.
  • MSHG-MAE consistently outperformed baseline models on multiple regression tasks with significant RMSE improvements.
  • Achieved RMSEs of 0.465 for ESOL, 0.780 for FreeSolv, and 0.501 for Lipophilicity, each showing considerable reductions compared to baseline methods.
  • Maintained structural and functional-group geometry without degradation in learned representations.

Abstract

Molecular representation learning often focuses on recognizing local structural patterns but struggles to capture functional groups and cross-scale interactions in a chemically meaningful way. Here, we propose the Multiscale Hypergraph Convolutional Masked Autoencoder (MSHG-MAE), a hypergraph-based self-supervised framework for drug-like molecular representation learning. We model each molecule as a unified molecular hypergraph with atom nodes and multitype hyperedges, including bond, ring, functional group, conjugated system, and hydrogen bond hyperedges, and use multiscale hypergraph convolutions to jointly capture dependencies at the atomic, substructural, and molecular levels. In pretraining, a semantics-aware masked autoencoding objective applies biased masking to atoms and hyperedges and reconstructs masked node features and hyperedge attributes, encouraging the model to internalize functional-group semantics. Furthermore, we introduce Δ-Property Alignment (Δ-PropAlign), which aligns embedding differences with proxy property differences so that the learned representations remain sensitive to interpretable structure–property changes. Experiments on multiple public molecular property benchmarks under scaffold splits show that MSHG-MAE consistently outperforms baselines on regression tasks and achieves competitive results on classification benchmarks. On three physicochemical regression benchmarks under scaffold splits (ESOL, FreeSolv, and Lipophilicity), MSHG-MAE with Δ-PropAlign achieves RMSEs of 0.465, 0.780, and 0.501, respectively, corresponding to approximately 40% and 47% lower RMSE than Uni-Mol on ESOL and FreeSolv and approximately 27% lower RMSE than D-MPNN on Lipophilicity (Table 5, scaffold split 80/10/10, 3 seeds), while not degrading structural or functional-group geometry. These results indicate that Δ-PropAlign improves the consistency between embedding differences and property differences without sacrificing structural or functional-group organization in the learned representations. The core code is publicly available on GitHub (https://github.com/Irzos/MSHG-MAE), and the data sets used in this study are publicly available from ZINC20 and MoleculeNet.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2026) studied this question.

synapsesocial.com/papers/69ba44654e9516ffd37a6126https://doi.org/10.1021/acs.jcim.5c02994
Ask AI
Helpful
Bookmark
Share
View Full Paper