PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 2026Journal of Chemical Information and Modeling3 citationsOpen Access

Large Language Model Agent for Modular Task Execution in Drug Discovery

View Full Paper
JOJanghoon OckRMRadheesh Sharma MedaSBSrivathsan Badrinarayanan

Key Points

  • The aim is to enhance the drug discovery pipeline by integrating large language models with modular tools for key tasks.
  • Developed a framework combining LLM reasoning with domain-specific tools.
  • Automated data retrieval including FASTA sequences and SMILES representations.
  • Performed multiproperty prediction and molecular refinement across iterations.
  • Generated 3D protein-ligand complexes and estimated binding affinities.
  • Molecules with QED >0.6 increased from 34 to 55 after two refinement rounds.
  • Compliance with the Ghose filter improved from 32 to 55 within 100 molecules.
  • Achieved 75 predicted properties related to ADMET and physicochemical descriptors.

Abstract

We present a modular framework powered by large language models (LLMs) that automates and streamlines key tasks across the early stage computational drug discovery pipeline. By combining LLM reasoning with domain-specific tools, the framework performs biomedical data retrieval, literature-grounded question answering via retrieval-augmented generation, molecular generation, multiproperty prediction, property-aware molecular refinement, and 3D protein-ligand structure generation. The agent autonomously retrieves relevant biomolecular information, including FASTA sequences, SMILES representations, and literature, and answers mechanistic questions with improved contextual accuracy compared to standard LLMs. It then generates chemically diverse seed molecules and predicted 75 properties, including ADMET-related and general physicochemical descriptors, which guids iterative molecular refinement. Across two refinement rounds, the number of molecules with QED >0.6 increased from 34 to 55. The number of molecules satisfying empirical drug-likeness filters also rose; for example, compliance with the Ghose filter increased from 32 to 55 within a pool of 100 molecules. The framework also employed Boltz-2 to generate 3D protein-ligand complexes and provide rapid binding affinity estimates for candidate compounds. These results demonstrate that the approach effectively supports molecular screening, prioritization, and structure evaluation. Its modular design enables flexible integration of evolving tools and models, providing a scalable foundation for AI-assisted therapeutic discovery.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ock et al. (2026) studied this question.

synapsesocial.com/papers/698d6d695be6419ac0d52478https://doi.org/10.1021/acs.jcim.5c02454
Ask AI
Helpful
Bookmark
Share
View Full Paper