PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 21, 20252 citationsOpen Access

Protein structure alignment significance is often exaggerated

View Full Paper
RER. C. EdgarHSHarutyun Sahakyan

Key Points

  • Main finding shows that unrelated proteins commonly exhibit convergent evolution, leading to high false positive rates.
  • Key evidence indicates previous alignment algorithms overestimate significance up to six orders of magnitude, impacting research accuracy.
  • Approach includes a novel method for estimating statistical significance that scales effectively with database size in protein searches.
  • Significance lies in providing more accurate statistical measures, advancing the reliability of computational protein structure analysis.

Abstract

Machine learning has generated millions of high-quality predicted protein structures, creating a need for computationally efficient structure search algorithms and robust estimates of statistical significance at this scale. We show that unrelated proteins have a universal tendency towards convergent evolution of secondary and tertiary motifs, causing an excess of high-scoring false positive alignments. To address this excess, and to accommodate recent innovations in search algorithm design, we describe a novel method for estimating statistical significance. We implement our approach in Reseek, showing that its E -values are accurate, scale successfully with database size, and are robust against the (generally unknown) diversity of folds in the database. We investigate popular structure search and alignment algorithms, finding that previous methods routinely overestimate significance by up to six orders of magnitude.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Edgar et al. (2025) studied this question.

synapsesocial.com/papers/689a02c9e6551bb0af8cd01chttps://doi.org/10.1101/2025.07.17.665375
Ask AI
Helpful
Bookmark
Share
View Full Paper