PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 6, 2026PeerJ Computer Science0 citationsOpen Access

The shift towards auto-labeling in natural language processing: a literature review on large language model-based methods

View Full Paper
HNHazem A. M. NabilSLSyaheerah Lebai LutfiHMHala Mulki

Key Points

  • This review aims to examine the role of large language models in automatic data annotation for natural language processing tasks.
  • Analyzed 38 peer-reviewed articles and preprints published between 2019 and 2025
  • Proposed a taxonomy categorizing LLM roles in annotation pipelines
  • Compared LLM-based auto-labeling to traditional human labeling methods.
  • LLMs approach but do not exceed human-level accuracy in auto-labeling tasks
  • Identified four key LLM roles in auto-labeling: DirectLabeler, PseudoLabeler, AssistantLabeler, DataGenerator
  • LLMs offer significant speed and cost advantages over manual labeling methods.

Abstract

Manual data annotation is a challenging and labor-intensive task for subjective natural language processing (NLP) applications like healthcare, cyberbullying, and sentiment analysis, where contextual understanding is required. Annotated data is essential for training supervised Machine Learning (ML) models. With the growing demand for high-quality labeled datasets, Large Language Models (LLMs) are increasingly used for automatic data annotation (auto-labeling). This review aims to explore the use of LLMs in auto-labeling NLP tasks, propose a taxonomy categorizing LLM roles in annotation pipelines, highlight pressing challenges, and provide practical insights and recommendations for future research. We analyzed 38 peer-reviewed articles and preprints published between 2019 and 2025, focusing on LLM applications in auto-labeling tasks such as sentiment analysis and stance detection. The review examines model types, training and prompting strategies, and compares LLM-based auto-labeling to traditional human labeling methods. We introduce a taxonomy that identifies four key LLM roles in auto-labeling: DirectLabeler, PseudoLabeler, AssistantLabeler, and DataGenerator. Across different fine-tuning and prompting techniques, findings show that LLMs generally approach and do not surpass human-level accuracy, while offering significant speed and cost advantages. Closed-source models dominate by frequency, yet open-source alternatives like Bidirectional Encoder Representations from Transformers (BERT) model remain important. Zero-shot and few-shot prompting techniques are common due to their convenience, though fine-tuning yields better domain-specific results. Challenges include prompt sensitivity, limited domain generalization, and ethical considerations. The proposed taxonomy offers a structured framework to understand and develop LLM-based auto-labeling methods, with scope for future expansion as the field evolves. This highlighted the potential of LLMs to automate and enhance data labeling across NLP tasks. Based on the identified strengths and limitations, our recommendations emphasize domain adaptation, fine-tuning on diverse datasets, iterative prompt engineering, and privacy-conscious open-source usage. Our findings suggest that with continued advancements in research and careful deployment, LLMs are poised to improve annotation scalability and efficiency in NLP and beyond.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nabil et al. (2026) studied this question.

synapsesocial.com/papers/6a23bb9a71a5da9775e77088https://doi.org/10.7717/peerj-cs.3878
Ask AI
Helpful
Bookmark
Share
View Full Paper