The framework improves entity recognition accuracy in text by integrating LLMs and structured boundary labeling.
Named Entity Recognition (NER) aims to identify entities with specific semantic meanings from text and classify them into predefined categories. With the rapid progress of generative large language models (LLMs), their strong capabilities in text comprehension and generation have sparked growing interest in applying them to information extraction tasks, including NER, within a generative paradigm. However, LLMs are often regarded as "black-box" systems, making it difficult to interpret the reasoning behind their predictions in entity recognition. In contrast, traditional BERT-based models explicitly assign labels to each token, ensuring precise capture of entity boundaries, while LLMs tend to infer spans from context, which may weaken their sensitivity to boundary information.To overcome these challenges, we propose a novel framework: CBATD-PC. Our approach integrates an instruction-tuning strategy for model construction and introduces a character-level boundary labeling mechanism. By decoupling the NER process into two stages, boundary detection and type classification, we establish a more structured and interpretable prediction pipeline. Moreover, by leveraging the contextual reasoning strengths of LLMs, we design a post-correction module that fine-tunes and refines the initial extraction results, thereby improving the overall accuracy and robustness of entity recognition.
No takes yet. Share an insight, caveat, or question.
Yongshen et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: