PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 9, 2024International Journal of Power Electronics and Drive Systems/International Journal of Electrical and Computer Engineering3 citationsOpen Access

Named entity recognition on Indonesian legal documents: a dataset and study using transformer-based models

View Full Paper
EYEvi YuliantiNBNaradhipa BharyJAJafar Abdurrohman

Key Points

Key points are not available for this paper at this time.

Abstract

The large volume of court decision documents in Indonesia poses a challenge for researchers to assist legal practitioners in extracting useful information from the documents. This information can also benefit the general public by improving legal transparency, law enforcement, and people's understanding of the law implementation in Indonesia. A natural language processing task that extracts important information from a document is called named entity recognition (NER). In this study, the NER task is applied to legal domains, which is then referred to as legal entity recognition (LER) task. In this task, some important legal entities, such as judges, prosecutors, and advocates, are extracted from the decision documents. A new Indonesian LER dataset is built, called IndoLER data, consisting of approximately 1K decision documents with 20 types of fine-grained legal entities. Then, the transformer-based models, such as multilingual bidirectional encoder representations from transformers (BERT) or M-BERT, Indonesian BERT or IndoBERT, Indonesian robustly optimized BERT pretraining approach (RoBERTa) or IndoRoBERTa, XLM (cross lingual language model)-RoBERTa or XLMR, are proposed to solve the Indonesian LER task using this dataset. Our experimental results show that the RoBERTa-based models, such as XLM-R and IndoRoBERTa, can outperform the state-of-the-art deep-learning baselines using BiLSTM (bidirectional long short-term memory) and BiLSTM-conditional random field (BiLSTM-CRF) approaches by 7.2% to 7.9% and 2.1% to 2.6%, respectively. XLM-RoBERTa is shown to be the best-performing model, achieving the F1-score of 0.9295.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yulianti et al. (2024) studied this question.

synapsesocial.com/papers/68e5cdb1b6db643587563a3chttps://doi.org/10.11591/ijece.v14i5.pp5489-5501
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Named Entity Recognition On Legal Documents Using Legal Bert2026
  2. 2A workflow-oriented and risk-aware system for Turkish legal named entity recognition: integrating transformer-based models with legal knowledge graphs2026
  3. 3EVALUATION OF INDOBERT AND ROBERTA: PERFORMANCE OF INDONESIAN LANGUAGE TRANSFORMER MODELS IN SENTIMENT CLASSIFICATION2025 · 1 citations
  4. 4COMPARATIVE PERFORMANCE OF TRANSFORMER AND LSTM MODELS FOR INDONESIAN INFORMATION RETRIEVAL WITH INDOBERT2025
  5. 5Leveraging convolutional neural network and transformer synergy for robust legal entity recognition in Chinese documents2026