ABSTRACT Artificial intelligence (AI), such as large language models (LLMs), is increasingly being used to score psychological and organizational constructs from text. However, the reliability and validity of these scores remains vital. Frequently, AI scores are treated as replicants of human ratings and inter‐rater reliability (IRR) is established by correlating AI and subject‐matter‐expert (SME) ratings. Yet, because LLMs may meaningfully differ from humans in ability to score text, correlating LLM scores with SME ratings is inappropriate because correlational approaches assume that AI and human scores are parallel, a condition that is unlikely and difficult to verify. This study proposes a factor‐analytic alternative in which AI scores and SME ratings are modeled as congeneric indicators of a common latent construct, estimating AI IRR as the squared standardized AI loading. Using a Monte Carlo simulation, correlational and CFA‐based IRR estimates were compared while manipulating true LLM and SME reliability, sample size, number of SME raters, and rater mean differences under ill‐structured rater designs. The simulations confirmed that when parallelism is violated, correlation‐based estimates (whether using single‐rater or composite SME ratings) are severely biased. In contrast, the proposed CFA framework accurately estimates AI IRR across diverse conditions, with precision improving as sample size, true IRR, and the number of SME raters increases. Furthermore, although ill‐structured designs reduce estimation accuracy, within‐rater standardization eliminated these inaccuracies. In conclusion, this paper urges researchers and practitioners to stop using correlational designs when estimating the IRR of LLM scores and offers best practice guidance going forward.
Andrew B. Speer (Thu,) studied this question.