PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 20, 2025Political Analysis20 citationsOpen Access

Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts

View Full Paper
AHAndrew HaltermanKKKatherine A. Keith

Key Points

  • Current open-weight large language models face challenges in accurately following codebook guidelines and measuring political constructs.
  • Using three curated political science codebooks, we demonstrate that supervised instruction-tuning significantly enhances LLM performance.
  • Our five-stage framework outlines how to prepare and evaluate LLMs against real-world codebook operationalizations for political analysis.
  • This contributes valuable datasets and guidance for researchers looking to implement codebook-LLM measurement projects effectively.

Abstract

Abstract Codebooks—documents that operationalize concepts and outline annotation procedures—are used almost universally by social scientists when coding political texts. To code these texts automatically, researchers are increasingly turning to generative large language models (LLMs). However, there is limited empirical evidence on whether “off-the-shelf” LLMs faithfully follow real-world codebook operationalizations and measure complex political constructs with sufficient accuracy. To address this, we gather and curate three real-world political science codebooks—covering protest events, political violence, and manifestos—along with their unstructured texts and human-coded labels. We also propose a five-stage framework for codebook-LLM measurement: Preparing a codebook for both humans and LLMs, testing LLMs’ basic capabilities on a codebook, evaluating zero-shot measurement accuracy (i.e., off-the-shelf performance), analyzing errors, and further (parameter-efficient) supervised training of LLMs. We provide an empirical demonstration of this framework using our three codebook datasets and several pre-trained 7–12 billion open-weight LLMs. We find current open-weight LLMs have limitations in following codebooks zero-shot, but that supervised instruction-tuning can substantially improve performance. Rather than suggesting the “best” LLM, our contribution lies in our codebook datasets, evaluation framework, and guidance for applied researchers who wish to implement their own codebook-LLM measurement projects.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Halterman et al. (2025) studied this question.

synapsesocial.com/papers/68d46aa631b076d99fa67467https://doi.org/10.1017/pan.2025.10017
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Large Language Models are Zero-Shot Reasoners2022 · 1,118 citations
  2. 2Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics2023 · 14,356 citations
  3. 3Stance detection: a practical guide to classifying political beliefs in text2024 · 28 citations
  4. 4The Measurement of Observer Agreement for Categorical Data1977 · 81,002 citations
  5. 5Measuring political violence in Pakistan: Insights from the BFRS Dataset2014 · 61 citations