Key points are not available for this paper at this time.
Abstract: Large language models (LLMs) have gained increased attention as tools for automatic item and scale generation in psychological assessment. Yet, it is not known how prompt instructions can affect the psychometric quality of LLM-generated items and scales. To test whether different prompt instructions influence the quality of generated items and scales, a general-purpose LLM (GPT4o) was prompted with seven different prompt instructions (one adhering to prompt engineering principles solely and six combining these principles with five different psychometric characteristics). Items were generated for two constructs, one established (emotional stability) and one rather novel (climate change anxiety). The LLM-generated items and scales were evaluated in terms of content validity, item difficulty, model fit, factor structure, reliability, and validity. Overall, there were no systematic differences regarding content validity, item difficulty, model fit, reliability and convergent validity, but unsystematic differences in terms of validity.
Oeljeklaus et al. (Wed,) studied this question.