PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 11, 2026PLoS Computational Biology0 citationsOpen Access

Trainable subnetworks reveal insights into structure knowledge organization in protein language models

RVRia VinodAAAva P. AminiLCLorin Crawford

Key Points

  • The research aims to explore how protein language models organize structural knowledge through trainable subnetworks.
  • Introduced trainable subnetworks that mask PLM weights related to language modeling for specific protein categories.
  • Trained 39 PLM subnetworks focusing on sequence- and residue-level features at different resolutions.
  • Used CATH taxonomy and secondary structure annotations to define structural categories.”
  • PLMs show high sensitivity to sequence-level features influencing structure prediction.
  • Factorized PLM representations respond significantly to small performance changes in language modeling, affecting predictions.
  • Disentangling coarse and fine-grained structural information reveals important insights into feature interactions.

Abstract

Protein language models (PLMs) pretrained via a masked language modeling objective have proven effective across a range of structure-related tasks, including high-resolution structure prediction. However, it remains unclear to what extent these models factorize protein structural categories among their learned parameters. In this work, we introduce trainable subnetworks, which mask out the PLM weights responsible for language modeling performance on a structural category of proteins. We systematically trained 39 PLM subnetworks targeting both sequence- and residue-level features at varying degrees of resolution using annotations defined by the CATH taxonomy and secondary structure elements. Using these PLM subnetworks, we assessed how structural factorization in PLMs influences downstream structure prediction. Our results show that PLMs are highly sensitive to sequence-level features and can predominantly disentangle extremely coarse or fine-grained information. Furthermore, we observe that structure prediction is highly responsive to factorized PLM representations and that small changes in language modeling performance can significantly impair PLM-based structure prediction capabilities. Our work presents a framework for studying feature entanglement within pretrained PLMs and can be leveraged to improve the alignment of learned PLM representations with known biological concepts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Vinod et al. (2026) studied this question.

synapsesocial.com/papers/698c1cb3267fb587c655f428https://doi.org/10.1371/journal.pcbi.1013925
Ask AI
Helpful
Bookmark
Share
View Full Paper