PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 3, 2026Behavior Research Methods1 citationsOpen Access

Generative psychometrics via AI-GENIE: Automatic item generation and validation with network-integrated evaluation

LRLara L Russell-LasalandraACAlexander P. ChristensenHGHudson Golino

Key Points

  • This study aims to present a new methodology for generating and validating psychological assessment items using AI and network psychometrics.
  • Developed AI-GENIE for automatic item generation using large language models.
  • Conducted Monte Carlo simulations with multiple AI models to create item pools.
  • Empirically tested item effectiveness across five representative U.S. samples (N=4,964).
  • AI-GENIE-generated scales achieved structural validity similar to expert-developed measures.
  • Notable improvements in item selection efficiency with increases of 8.68–20.03 in normalized mutual information.
  • Demonstrated application of AI-GENIE for measuring emerging constructs like AI anxiety.

Abstract

Abstract The rapid advancement of artificial intelligence (AI), particularly large language models (LLMs), has introduced powerful tools for various research domains, including psychological scale development. This study presents a methodology for efficiently generating and selecting high-quality, non-redundant items for psychological assessments using LLMs and network psychometrics. Our approach, termed Automatic Item Generation and Validation with Network-Integrated Evaluation (AI-GENIE), reduces reliance on expert intervention by integrating generative AI with the latest network psychometric techniques. The efficacy of AI-GENIE was evaluated through Monte Carlo simulations using the Mixtral, Gemma 2, Llama 3, GPT-3. 5, and GPT-4o models to generate item pools that mimic Big Five personality assessments. Additionally, items from AI-GENIE were empirically tested with five nationally representative U. S. samples (N = 4, 964 N = 4, 964 total), demonstrating that AI-GENIE-generated scales achieve structural validity—that is, evidence based on internal structure (dimensionality and item stability) —comparable to traditional expert-developed measures. The results demonstrated improvements in item selection efficiency, with overall average increases of 8. 68–20. 03 in normalized mutual information in the final item pool across all models. We also present a simulation study on the emerging construct of AI anxiety to demonstrate AI-GENIE’s utility for underrepresented constructs. Results from newly released models (DeepSeek, GPT-OSS 20B, GPT-OSS 120B) are presented in the Appendix. The findings suggest that AI-GENIE can streamline the scale development and structural validation process.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Russell-Lasalandra et al. (2026) studied this question.

synapsesocial.com/papers/6a4752405c29257aa25790d4https://doi.org/10.3758/s13428-026-03082-1
Ask AI
Helpful
Bookmark
Share
View Full Paper