Topic Modeling approaches face difficulties in processing legal texts because of their unique characteristics, such as the length of the texts and the specialized terminology used within them. The process of topic modeling involves finding a text's semantic structure. This way, specific approaches are needed. When the legal documents are presented has a lot to do with what topics are important. This paper aims to explain and evaluate BERTopic's application to topic modeling in legal documents. In this research, we experiment with BERTopic by utilizing its several pre-trained Arabic language models as embeddings. Performance evaluation employs the Normalized Pointwise Mutual Information (NPMI) measure. Notably, in comparison to multilingual pre-trained models, our findings reveal that BERTopic using Arabic monolingual pre-trained models exhibits superior performance, offering insights into sustainable and efficient topic modeling for legal documents.
No takes yet. Share an insight, caveat, or question.
Aouichaty et al. (2024) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: