PulseTrendingJournal ClubResearchersJournalsExplore
Instagram
HomeTrendingJournal ClubExplore
Synapse
⌘+K
Synapse
August 22, 2026ACM Transactions on Information Systems

GenTREC : The First Test Collection Generated by Large Language Models for Evaluating Information Retrieval Systems

View Full Paper
Ask AI
Bookmark
Share

Authors

MTMehmet Deniz TürkmenMKMücahid KutluBABahadır ALTUN

Discussion

Loading...

Member takes

Overview

Evaluation study demonstrates that LLM-generated documents accurately rank information retrieval systems, indicating a low-cost alternative to manual test collections.

Key Points

  • To introduce and evaluate GenTREC, the first Information Retrieval test collection generated entirely by large language models to eliminate the need for manual relevance judgments.
  • Generated 96,196 synthetic documents (comprising both prompt-relevant and non-relevant materials) across 300 TREC search topics using a large language model.
  • Evaluated document quality, relevance judgment accuracy, and benchmarked Information Retrieval system rankings across P@100, MAP, RPrec, and nDCG metrics against traditional TREC collections.
  • Information Retrieval system rankings evaluated on GenTREC aligned consistently with traditional TREC test collections across P@100, MAP, RPrec, and nDCG metrics.
  • The synthetic dataset successfully eliminated manual annotation requirements while preserving the evaluation reliability needed to distinguish search system performance.

Cite This Study

Türkmen et al. (2026) studied this question.

synapsesocial.com/papers/6a895ec3ca7ade938187ce54https://doi.org/10.1145/3841465
View Full Paper
Ask AI
Bookmark
Share