PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 18, 2025Applied Sciences0 citationsOpen Access

Information Extraction from Multi-Domain Scientific Documents: Methods and Insights

View Full Paper
TBTatiana BaturaAYAigerim YerimbetovaНМНуржан Мукажанов

Key Points

  • The evaluation demonstrated high performance in entity recognition using spaCy and GLiNER, enhancing accessibility in low-resource languages.
  • Notable models like BERT, LLaMA, and spaCy were assessed for entity recognition in diverse domains such as IT and Medicine across Kazakh and Russian.
  • A novel zero-shot relation extraction model was introduced, enabling predictions of relations between entities in previously unseen documents.
  • Findings emphasize the growing need for targeted tools and resources in the under-resourced language segments, especially Kazakh and Russian.

Abstract

The rapid growth of scientific literature necessitates effective information extraction. However, existing methods face significant challenges, particularly when applied to multi-domain documents and low-resource languages. For Kazakh and Russian, there is a notable lack of annotated corpora and dedicated tools for scientific information extraction. To address this gap, we introduce SciMDIX (Scientific Multi-Domain Information extraction), a novel multi-domain dataset of scientific documents in Russian and Kazakh, annotated with entities and relations. Our study includes a comprehensive evaluation of entity recognition performance, comparing state-of-the-art models, such as BERT, LLaMA, GLiNER, and spaCy across four diverse domains (IT, Linguistics, Medicine, and Psychology) in both languages. The findings highlight the promise of spaCy and GLiNER for practical deployment in under-resourced language settings. Furthermore, we propose a new zero-shot relation extraction model that leverages a multimodal representation by integrating sentence context, entity mentions, and textual definitions of relation classes. Our model can predict semantic relations between entities in new documents, even for a language encountered during training. This capability is especially valuable for low-resource language scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Batura et al. (2025) studied this question.

synapsesocial.com/papers/68af4953ad7bf08b1ead50e2https://doi.org/10.3390/app15169086
Ask AI
Helpful
Bookmark
Share
View Full Paper