Evaluation study demonstrates that LLM-generated documents accurately rank information retrieval systems, indicating a low-cost alternative to manual test collections.
Key Points
To introduce and evaluate GenTREC, the first Information Retrieval test collection generated entirely by large language models to eliminate the need for manual relevance judgments.
Generated 96,196 synthetic documents (comprising both prompt-relevant and non-relevant materials) across 300 TREC search topics using a large language model.
Evaluated document quality, relevance judgment accuracy, and benchmarked Information Retrieval system rankings across P@100, MAP, RPrec, and nDCG metrics against traditional TREC collections.
Information Retrieval system rankings evaluated on GenTREC aligned consistently with traditional TREC test collections across P@100, MAP, RPrec, and nDCG metrics.
The synthetic dataset successfully eliminated manual annotation requirements while preserving the evaluation reliability needed to distinguish search system performance.