PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 19, 2026Open Linguistics0 citationsOpen Access

Large language models as first-pass filters for corpus annotation: semantic disambiguation of Galician pobo

View Full Paper
VMVítor Míguez-RegoUniversidade de Santiago de Compostela

Key Points

  • The research aims to evaluate the effectiveness of large language models as first-pass filters in semantic annotation, particularly for polysemous terms in low-resource languages.
  • Annotated 300 examples of the Galician noun 'pobo' by three human coders and four LLMs.
  • Employed a static, single-phase prompting approach for LLM annotation.
  • Prioritized recall over precision to capture occurrences of the target phenomenon.
  • Achieved F 2 score of 0.944 with Claude 4.5 Opus against human consensus.
  • Demonstrated 100% recall with no information loss in filtering tasks.
  • Showed substantial workload reduction when using LLMs for annotation.

Abstract

Abstract This paper demonstrates the use of LLMs as first-pass filters in corpus annotation, with a focus on semantic disambiguation – a task more challenging than form-based classification due to its context-dependence. Using as a case study the polysemous Galician noun pobo ‘people/village’, the study demonstrates the applicability of LLM-assisted annotation to low-resource languages. 300 examples were annotated by three human coders and four LLMs (Claude 4 Sonnet, Claude 4 Opus, Claude 4.5 Sonnet, and Claude 4.5 Opus) using a static, single-phase prompting approach. Since first-pass filters should capture as many actual occurrences of the target phenomenon as possible, priority was given to recall over precision. Accordingly, the paper argues for F 2 , a recall-focused metric, over commonly used alternatives like F 1 or MCC for validating LLM performance in filtering tasks. Claude 4.5 Opus with pretraining achieved the best performance against the human consensus ( F 2 = 0.944, recall = 100 %), resulting in substantial workload reduction with no information loss. The study demonstrates that LLMs can serve as effective first-pass filters for semantic annotation in corpus linguistics, extending their applicability to low-resource languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Vítor Míguez-Rego (2026) studied this question.

synapsesocial.com/papers/6996a7b5ecb39a600b3edb1fhttps://doi.org/10.1515/opli-2025-0078
Ask AI
Helpful
Bookmark
Share
View Full Paper