PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 24, 20246 citationsOpen Access

Beware of Data Leakage from Protein LLM Pretraining

View Full Paper
LHLeon HermannTFTobias FiedlerHNHoang An Nguyen

Key Points

Key points are not available for this paper at this time.

Abstract

Abstract Pretrained protein language models are becoming increasingly popular as a backbone for protein property inference tasks such as structure prediction or function annotation, accelerating biological research. However, related research oftentimes does not consider the effects of data leakage from pretraining on the actual downstream task, resulting in potentially unrealistic performance estimates. Reported generalization might not necessarily be reproducible for proteins highly dissimilar from the pretraining set. In this work, we measure the effects of data leakage from protein language model pretraining in the domain of protein thermostability prediction. Specifically, we compare two different dataset split strategies: a pretraining-aware split, designed to avoid similarity between pretraining data and the held-out test sets, and a commonly-used naive split, relying on clustering the training data for a downstream task without taking the pretraining data into account. Our experiments suggest that data leakage from language model pretraining shows consistent effects on melting point prediction across all experiments, distorting the measured performance. The source code and our dataset splits are available at https://github.com/tfiedlerdev/pretraining-aware-hotprot .

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hermann et al. (2024) studied this question.

synapsesocial.com/papers/68e5f2e3b6db643587587b1dhttps://doi.org/10.1101/2024.07.23.604678
Ask AI
Helpful
Bookmark
Share
View Full Paper