PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 20, 20250 citationsOpen Access

FreeRet: MLLMs as Training-Free Retrievers

View Full Paper
YZYuhan ZhuXZXiangyu ZengCWChao Wang

Key Points

  • FreeRet enables retrieval using pretrained MLLMs, showing that they can perform without additional training.
  • The framework achieved significant improvements over models trained on millions of pairs on 46 datasets.
  • It introduces three advances: semantically grounded embeddings, explicit prior conditioning, and neutral choice framing.
  • FreeRet is model-agnostic, maintaining generative abilities while integrating retrieval and generation processes.

Abstract

Multimodal large language models (MLLMs) are emerging as versatile foundations for mixed-modality retrieval. Yet, they often require heavy post-hoc training to convert them into contrastive encoders for retrieval. This work asks: Can off-the-shelf MLLMs serve as powerful retrievers without additional training? We present FreeRet, a plug-and-play framework that turns any MLLM into a two-stage retriever. FreeRet first derives semantically grounded embeddings directly from the model for fast candidate search, and then exploits its reasoning ability for precise reranking. The framework contributes three advances: bypassing lexical alignment layers to obtain semantically faithful embeddings, conditioning representation generation with explicit priors, and mitigating framing effect in reranking via neutral choice framing. On the MMEB and MMEB-V2 benchmarks spanning 46 datasets, FreeRet substantially outperforms models trained on millions of pairs. Beyond benchmarks, FreeRet is model-agnostic and scales seamlessly across MLLM families and sizes, preserves their generative abilities, supports arbitrary modality combinations, and unifies retrieval, reranking, and generation into end-to-end RAG within a single model. Our findings demonstrate that pretrained MLLMs, when carefully harnessed, can serve as strong retrieval engines without training, closing a critical gap in their role as generalists.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2025) studied this question.

synapsesocial.com/papers/68f5fcce8d54a28a75cf1c97https://doi.org/10.48550/arxiv.2509.24621
Ask AI
Helpful
Bookmark
Share
View Full Paper