The dataset enhances machine reading comprehension in Urdu by providing 20,000 annotated pairs and a new evaluation metric, semantic match.
Machine Reading Comprehension (MRC) is a key task in Natural Language Understanding that enables automated systems to answer questions based on textual input. While MRC has made substantial progress for high-resource languages, low-resource languages pose significant challenges due to their complex linguistic features. This paper presents a comprehensive human-annotated dataset for Urdu MRC, comprising 20,000 question-answer pairs derived from 1,540 articles across seven domains. Unlike previous translation-based datasets, this dataset contains question-answer pairs created through rigorous crowd-sourcing and expert annotation. The dataset encompasses diverse question types, including both answerable and unanswerable questions, with answers ranging from single words to complete sentences, effectively capturing Urdu’s morphological richness and syntactic diversity. To address the limitations of traditional evaluation metrics like Exact Match (EM) and F 1 in assessing Urdu answers, we propose Semantic Match (SM), a metric designed to measure semantic equivalence between predicted and ground-truth answers. Our evaluation demonstrates the dataset’s increased complexity, with state-of-the-art models achieving only 0.82% SM accuracy. Together, the dataset and evaluation metric establish a robust framework for advancing Urdu MRC research, bridging critical gaps in both dataset quality and evaluation methodology.
No takes yet. Share an insight, caveat, or question.
Kazi et al. (2025) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: