This study investigated a simpler alternative to cumbersome multisensor systems (e.g., multi-channel EMG) for articulatory gesture recognition. The method was validated in a vowel classification task analyzing pulse-echo waveforms from articulatory muscles. Pulse-echo signals for five Japanese vowels and a neutral state were acquired from two male subjects (20s). Two 5-MHz ultrasonic transducers (UTs) with a diameter of 10 mm were placed on the chin near the digastric muscle. Signals were acquired using two pulser-receivers (pulse repetition frequency: 2 kHz) and an oscilloscope (sampling frequency: 625 MHz). Five 6-s trials (2 s neutral, 2 s vowel, 2 s neutral) were conducted for each vowel. Spatial features were obtained using a discrete wavelet transform and mean absolute value, while temporal features were derived from their time differences. After feature selection, k-nearest neighbors (kNN) and linear discriminant analysis (LDA) classifiers were trained. Performance was validated using trial-based 5-fold cross-validation. For comparison, the same experiments and analyses were conducted with UTs on both cheeks near the masseter muscles. The LDA classifier achieved higher accuracies of 94% (digastric muscle) and 83% (masseter muscle) compared to the kNN classifier with 80% and 70%, respectively. Similar results from the second subject support the method's effectiveness.
Kanaya et al. (Wed,) studied this question.