Abstract Survey researchers rely heavily on closed-ended questions to measure latent respondent characteristics like knowledge, policy positions, emotions, ideology, and various other traits. Closed-ended questions are easy to analyze and collect, but necessarily limit the depth and variability of responses. Open-ended responses allow for greater depth and variability in responses, but are labor intensive to code. Large language models (LLMs) may help with this problem, but existing approaches to using LLMs have a number of limitations. In this paper, we propose and test a pairwise comparison method to scale open-ended survey responses on a continuous scale. The approach relies on LLMs to make pairwise comparisons of statements that identify which statement “wins” and “loses.” With this information, we employ a Bayesian Bradley-Terry model to recover a “score” on a latent dimension for each statement. This approach allows for finer discrimination between items, reduced anchoring bias, better measurement of uncertainty, and is more flexible than methods relying on Maximum Likelihood Estimation techniques. We demonstrate the utility of this approach on an open-ended question probing knowledge of interest rates in the US economy. A comparison of six LLMs of various sizes reveals that pairwise comparisons show greater consistency than zero-shot 0–10 ratings across a variety of model sizes. Further, comparison of pairwise decisions is consistent with knowledgeable crowdsourced workers.
DiGiuseppe et al. (2026) studied this question.