PulseTrendingJournal ClubResearchersJournalsExplore
Instagram
HomeTrendingJournal ClubExplore
Synapse
⌘+K
Synapse
June 20, 2026Data in BriefOpen Access

KannadaLit4NLP: A Comprehensive Classical Kannada Literary Dataset of Vachanas, Tripadis, and Kagga with Scholarly Interpretations for Natural Language Processing

View Full Paper
Ask AI
Bookmark
Share

Authors

BCBasavanna CSMS. ManjunathDGD. S. Guru

Discussion

Loading...

Member takes

Overview

Dataset provides Kannada literary texts and interpretations for NLP research, enhancing language representation.

Key Points

  • To develop a comprehensive corpus of classical Kannada literature to support NLP research for low-resource languages.
  • Created a dataset of 24,746 literary verses and 22,369 interpretations from Kannada traditions.
  • Used optical character recognition for digitization and manual verification for accuracy.
  • Organized entries with metadata for various NLP applications such as semantic understanding and translation.
  • The dataset integrates historical and modern Kannada literature with scholarly interpretations.
  • Supports various NLP tasks including semantic textual similarity and information retrieval.
  • Publicly available to promote further research and reproducibility in Kannada NLP.

Cite This Study

C et al. (2026) studied this question.

synapsesocial.com/papers/6a362fbcdb0793dc1a537285https://doi.org/10.1016/j.dib.2026.112983
View Full Paper
Ask AI
Bookmark
Share