PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 15, 2026International Journal of Information and Communication Technology0 citationsOpen Access

An English vocabulary pronunciation evaluation model based on multidimensional audio features and machine learning

CDCan Du

Key Points

  • The research aims to improve English vocabulary pronunciation evaluation by utilizing advanced audio features and machine learning techniques.
  • Designed a multi-dimensional audio feature extraction algorithm using multi-scale dilated convolution.
  • Constructed a shallow feature refinement module with parallel convolutions for capturing three-dimensional features.
  • Implemented a global feature fusion module with multiplicative gating mechanisms for cross-scale feature fusion.
  • Used a differential evolution algorithm optimized support vector machine to score multi-dimensional features.
  • Achieved an average evaluation accuracy of 94.57%.
  • Outperformed comparative models in pronunciation assessment accuracy.
  • Provided a more objective and accurate evaluation of English vocabulary pronunciation.

Abstract

In response to the issue where current English vocabulary pronunciation evaluation models cannot fully extract feature information from different dimensions of spectrograms, this paper first designs a multi-dimensional audio feature extraction algorithm based on multi-scale dilated convolution.This algorithm initially constructs a shallow feature refinement module that uses parallel convolutions to capture time, frequency, and time-frequency three-dimensional shallow features of Mel-frequency cepstral coefficients features.It combines Res2net structure, dilated convolution, and channel attention to capture more fine-grained multi-scale information from the shallow multi-dimensional features.Then it employs a global feature fusion module combined with multiplicative gating mechanisms to enhance cross-scale feature fusion.Finally, differential evolution algorithm optimised support vector machines are used to score the multi-dimensional features.Experimental results indicate that the average evaluation accuracy of the proposed model reaches 94.57%, outperforming comparative models and achieving an objective and accurate assessment of English vocabulary pronunciation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Can Du (2026) studied this question.

synapsesocial.com/papers/69b64ccdb42794e3e660dfabhttps://doi.org/10.1504/ijict.2026.152224
Ask AI
Helpful
Bookmark
Share
View Full Paper