PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 20250 citationsOpen Access

Language Model Re-rankers are Fooled by Lexical Similarities

View Full Paper
LHLovisa HagströmENErcong NieRHRuben Halifa

Key Points

  • Results indicate that LM re-rankers struggled to surpass a simple BM25 baseline on the DRUID dataset, revealing performance issues.
  • Evaluation revealed that the performance of LM re-rankers mainly improved on the NQ dataset, pointing to dataset-specific strengths.
  • The study identifies weaknesses in LM re-rankers through analysis of errors related to lexical dissimilarities during retrieval.
  • Findings emphasize the necessity for more adversarial datasets to robustly evaluate language model re-rankers’ true capabilities.

Abstract

Language model (LM) re-rankers are used to refine retrieval results for retrieval-augmented generation (RAG). They are more expensive than lexical matching methods like BM25 but assumed to better process semantic information and the relations between the query and the retrieved answers. To understand whether LM re-rankers always live up to this assumption, we evaluate 6 different LM re-rankers on the NQ, LitQA2 and DRUID datasets. Our results show that LM re-rankers struggle to outperform a simple BM25 baseline on DRUID. Leveraging a novel separation metric based on BM25 scores, we explain and identify re-ranker errors stemming from lexical dissimilarities. We also investigate different methods to improve LM re-ranker performance and find these methods mainly useful for NQ. Taken together, our work identifies and explains weaknesses of LM re-rankers and points to the need for more adversarial and realistic datasets for their evaluation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hagström et al. (2025) studied this question.

synapsesocial.com/papers/68de84bb5b556a9128e1baechttps://doi.org/10.48550/arxiv.2502.17036
Ask AI
Helpful
Bookmark
Share
View Full Paper