Benchmarking study demonstrates high accuracy of fine-tuned language models in criminal court rulings, highlighting scalable legal information extraction.
General-purpose Named Entity Recognition (NER) models struggle with domain-specific challenges, particularly in the legal sector, where complex syntax, specialized terminology, and privacy concerns pose significant obstacles. This paper presents a novel data-driven two-stage framework that leverages Large Language Models (LLMs) fine-tuned for legal applications to enhance NER for criminal process documents. Our contributions include: (i) the design of a novel framework for structured entity extraction in legal texts, (ii) the definition of an enhanced ontology adapted to criminal law, and (iii) a comprehensive evaluation of the framework on a real-world dataset. As a side contribution of the paper, we extend the renowned OntoNotes5 ontology by integrating new domain-specific entities tailored to criminal law. To evaluate our approach, we construct and manually annotate a real-world dataset comprising over 100 criminal judgments from four Italian courts, sourced from the Italian Antimafia and Anti-terrorism National Directorate (Direzione Nazionale Antimafia e Antiterrorismo — DNAA). Experimental results demonstrate the effectiveness of fine-tuned LLMs in accurately identifying legal entities. Among the evaluated models, LLaMA-3.2-1B demonstrates the lowest training time, while LLaMA-2.7B achieves the fastest inference time. In terms of predictive performance, LLaMA-2.7B and Vicuna-7B consistently yield the highest accuracy scores across evaluation metrics.
No takes yet. Share an insight, caveat, or question.
Romano et al. (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: