PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 27, 20240 citationsOpen Access

Voices Unheard: NLP Resources and Models for Yor\`ub\'a Regional Dialects

View Full Paper
OAOrevaoghene AhiaAAAnuoluwapo AremuDADiana Abagyan

Key Points

Key points are not available for this paper at this time.

Abstract

Yor\`ub\'a an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in disparities for dialects and varieties for which there are little to no resources or tools. We take steps towards bridging this gap by introducing a new high-quality parallel text and speech corpus YOR\`ULECT across three domains and four regional Yor\`ub\'a dialects. To develop this corpus, we engaged native speakers, travelling to communities where these dialects are spoken, to collect text and speech data. Using our newly created corpus, we conducted extensive experiments on (text) machine translation, automatic speech recognition, and speech-to-text translation. Our results reveal substantial performance disparities between standard Yor\`ub\'a and the other dialects across all tasks. However, we also show that with dialect-adaptive finetuning, we are able to narrow this gap. We believe our dataset and experimental analysis will contribute greatly to developing NLP tools for Yor\`ub\'a and its dialects, and potentially for other African languages, by improving our understanding of existing challenges and offering a high-quality dataset for further development. We release YOR\`ULECT dataset and models publicly under an open license.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ahia et al. (2024) studied this question.

synapsesocial.com/papers/68e6312bb6db6435875c3aa4https://doi.org/10.48550/arxiv.2406.19564
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages2025
  2. 2Automatic Diacritization Models for a High-Population Low-Resource African Language (Yorùbá)2026
  3. 3BENYO-S2ST-Corpus-1: A Bilingual English-to-Yoruba Direct Speech-to-Speech Translation Corpus2025
  4. 4Building Text and Speech Benchmark Datasets and Models for Low‐Resourced East African Languages: Experiences and Lessons2024 · 11 citations
  5. 5Improving Tone Recognition Performance using Wav2vec 2.0-Based Learned Representation in Yoruba, a Low-Resourced Language2024 · 3 citations