Los puntos clave no están disponibles para este artículo en este momento.
The process of obtaining reference patterns for syllablelike units is tedious and time-consuming. As such, speech recognition systems based on such units are usually tested on only a single talker. In this talk we describe a procedure for using demisyllable reference patterns excised from spoken utterances for one talker, and automatically creating demisyllable reference patterns for a new talker. The procedure is based on dynamic time-warping alignment of the spoken utterances, and the assumption that the optimum warping path identifies the best matching demisyllable within the utterance. The procedure has been used to create demisyllable reference patterns for two new talkers. Recognition tests were performed using the automatically created references with both a 100- and 1109-word vocabulary. Word accuracies greater than 90% were obtained for both talkers on the 100-word vocabulary. Using the 1109-word vocabulary, it was shown that the recognition accuracy using the automatically extracted demisyllables was only slightly worse than the recognition accuracy from reference patterns based on hand corrections applied to the demisyllables.
Rabiner et al. (Sun,) studied this question.