PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 10, 2025ACM Transactions on Intelligent Systems and Technology18 citations

Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review

View Full Paper
ZRZhyar Rzgar K RostamGKGábor Kertész

Key Points

  • Utilizing pre-trained language models improves accuracy in text classification tasks across various domains.
  • Review of 41 articles revealed the effectiveness of transformer-based models like BERT for domain-specific applications.
  • This systematic review followed rigorous criteria as per PRISMA guidelines, ensuring methodological integrity in findings.
  • Challenges in text classification stem from unique vocabulary and data imbalances in specialized domains, necessitating tailored approaches.

Abstract

The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural Language Processing (NLP) plays a crucial role in addressing this challenge, particularly in text classification tasks. While Large Language Models (LLMs) have achieved remarkable success in NLP, their accuracy can suffer in domain-specific contexts due to specialized vocabulary, unique grammatical structures, and imbalanced data distributions. In this Systematic Literature Review (SLR), we investigate the utilization of Pre-trained Language Models (PLMs) for domain-specific text classification. We systematically review 41 articles published between 2018 and January 2024, adhering to the PRISMA statement (Preferred Reporting Items for Systematic Reviews and Meta-Analyses). This review methodology involved rigorous inclusion criteria and a multi-step selection process employing AI-powered tools. We delve into the evolution of text classification techniques and differentiate between traditional and modern approaches. We emphasize transformer-based models and explore the challenges and considerations associated with using LLMs for domain-specific text classification. Furthermore, we categorize existing research based on various PLMs and propose a taxonomy of techniques used in the field. To validate our findings, we conducted a comparative experiment involving BERT, SciBERT, and BioBERT in biomedical sentence classification. Finally, we present a comparative study on the performance of LLMs in text classification tasks across different domains. In addition, we examine recent advancements in PLMs for domain-specific text classification and offer insights into future directions and limitations in this rapidly evolving domain.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rostam et al. (2025) studied this question.

synapsesocial.com/papers/68c18f409b7b07f3a0615feehttps://doi.org/10.1145/3763002
Ask AI
Helpful
Bookmark
Share
View Full Paper