PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 1, 201418 citations

Machine learning approaches to improving pronunciation error detection on an imbalanced corpus

View Full Paper
XYXuesong YangALAnastassia LoukinaKEKeelan Evanini

Key Points

Key points are not available for this paper at this time.

Abstract

In this paper, we investigate the task of phone-level pronunciation error detection as a binary classification problem, the performance of which is heavily affected by the imbalanced distribution of the classes in a manually annotated data set of non-native English. In order to address problems caused by this extreme class imbalance, methods for cost-sensitive learning (weighting inversely proportional to class frequencies) and over-sampling of synthetic instances (SMOTE) are investigated in order to improve classification performance. Experiments using classifiers consisting of features based on acoustic phonetics and word identity demonstrate that these machine learning approaches lead to performance improvements over the baseline system based on the extremely imbalanced data. In addition, several different types of classifiers were compared. Finally, the paper analyzes the robustness of classifier performance across different phones.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2014) studied this question.

synapsesocial.com/papers/6a186a911ca866914fc99ca7https://doi.org/10.1109/slt.2014.7078591
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Speaker identification on the SCOTUS corpus2008 · 611 citations
  2. 2Large-scale characterization of Mandarin pronunciation errors made by native speakers of European languages2013 · 18 citations
  3. 3Automated speech scoring for non-native middle school students with multiple task types2013 · 44 citations
  4. 4ASR based pronunciation evaluation with automatically generated competing vocabulary and classifier fusion2009 · 21 citations
  5. 5Improvement of segmental mispronunciation detection with prior knowledge extracted from large L2 speech corpus2011 · 15 citations