PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 15, 20250 citationsOpen Access

Benchmarking Large Language Models for Polymer Property Predictions

View Full Paper
SGSonakshi GuptaAMAkhlak MahmoodSSShivank Shukla

Key Points

  • LLM-based methods closely approach traditional models in performance but generally underperform in predictive accuracy.
  • Finetuning of LLaMA-3 consistently yields better results than GPT-3.5, indicating the advantage of open-source architecture.
  • Single-task learning proves to be more effective than multi-task learning for LLMs in capturing polymer properties.
  • Analysis shows general purpose LLMs have limitations in representing complex chemo-structural information compared to domain-specific embeddings.

Abstract

Machine learning has revolutionized polymer science by enabling rapid property prediction and generative design. Large language models (LLMs) offer further opportunities in polymer informatics by simplifying workflows that traditionally rely on large labeled datasets, handcrafted representations, and complex feature engineering. LLMs leverage natural language inputs through transfer learning, eliminating the need for explicit fingerprinting and streamlining training. In this study, we finetune general purpose LLMs -- open-source LLaMA-3-8B and commercial GPT-3.5 -- on a curated dataset of 11,740 entries to predict key thermal properties: glass transition, melting, and decomposition temperatures. Using parameter-efficient fine-tuning and hyperparameter optimization, we benchmark these models against traditional fingerprinting-based approaches -- Polymer Genome, polyGNN, and polyBERT -- under single-task (ST) and multi-task (MT) learning. We find that while LLM-based methods approach traditional models in performance, they generally underperform in predictive accuracy and efficiency. LLaMA-3 consistently outperforms GPT-3.5, likely due to its tunable open-source architecture. Additionally, ST learning proves more effective than MT, as LLMs struggle to capture cross-property correlations, a key strength of traditional methods. Analysis of molecular embeddings reveals limitations of general purpose LLMs in representing nuanced chemo-structural information compared to handcrafted features and domain-specific embeddings. These findings provide insight into the interplay between molecular embeddings and natural language processing, guiding LLM selection for polymer informatics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gupta et al. (2025) studied this question.

synapsesocial.com/papers/68f02c7d616531447b5f94d7https://doi.org/10.48550/arxiv.2506.02129
Ask AI
Helpful
Bookmark
Share
View Full Paper