PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 1, 20240 citationsOpen Access

The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data

View Full Paper
GPGeorgios ParaskevopoulosCTChara TsoukalaAKAthanasios Katsamanis

Key Points

Key points are not available for this paper at this time.

Abstract

The development of speech technologies for languages with limited digital representation poses significant challenges, pri- marily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models. Recent research has underscored the potential of leveraging weak supervision to augment the pool of available data. In this study, we compile an 800-hour corpus of Modern Greek from podcasts and employ Whisper large-v3 to generate silver transcriptions. This corpus is utilized to fine-tune our models, aiming to assess the efficacy of this approach in enhancing ASR performance. Our analysis spans 16 distinct podcast domains, alongside evaluations on established datasets for Modern Greek. The findings indicate consistent WER improvements, correlating with increases in both data volume and model size. Our study confirms that assembling large, weakly supervised corpora serves as a cost-effective strategy for advancing speech technologies in under-resourced languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Paraskevopoulos et al. (2024) studied this question.

synapsesocial.com/papers/68e59e8eb6db643587538972https://doi.org/10.21437/interspeech.2024-1830
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data2024
  2. 2GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement2024 · 2 citations
  3. 3Speech Recognition for Greek Dialects: A Challenging Benchmark2024 · 2 citations
  4. 4EuroSpeech: A Multilingual Speech Corpus2025
  5. 5MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research2024