Key points are not available for this paper at this time.
The growing adoption of Large Language Models (LLMs) in software engineering has generated considerable interest in their potential for Requirements Engineering (RE) tasks. However, evaluating LLM performance in RE contexts presents a fundamental challenge: the absence of established ground truth, compounded by the free-form nature of LLM outputs that resist automated comparison. This paper proposes a general methodology for systematically evaluating LLMs in RE tasks in the absence of traditional ground truth with three complementary strategies: replacing traditional ground truth with literature-based knowledge extraction to create reference standards; decomposing complex prompts into discrete closed-form questions that enable quantitative assessment; and optionally employing synthetic data generation for controlled parameter variation. The methodology is conceived to generalise across diverse RE evaluation contexts, providing researchers and practitioners with a systematic approach for assessing LLM capabilities in RE tasks where traditional ground truth is unavailable. We illustrate the methodology through several examples. In addition, to validate the reliability of decomposing complex tasks into closed-form prompts, we conducted a comparative analysis to verify if the strategy provides a quantitatively reliable assessment.
Sabatucci et al. (Sat,) studied this question.