The Cranfield paradigm has been the dominant approach to evaluate information retrieval systems for decades, but—in its classical form—has clear limitations when it comes to conversational search systems, which synthesize unique outputs in a dynamic multi-turn interaction with the user. User simulation, i.e., the interaction of a computer program with a retrieval system instead of a human user to generate plausible conversations as a basis for evaluation, was proposed several years ago as a way to integrate the dynamics of conversational systems into an evaluation framework. Seen as a distant vision for years, the advent of large language models has propelled this idea forward. In 2025, there were the first three shared tasks in information retrieval where user simulation was used for evaluation or was the participants' goal. In this article, the organizers of these three shared tasks report on their specific evaluation approaches, highlight differences in setup, report on insights gained, and look to the future to discuss how user simulation can be integrated into a new evaluation paradigm.
Gohsen et al. (Mon,) studied this question.