Describes an approach to the automatic evaluation of both the speech recognition and understanding capabilities of a spoken dialogue system for train timetable information. For performance judgement, we use word accuracy for recognition and concept accuracy for understanding. Both measures are calculated by comparing these modules' outputs with a correct reference answer. We report evaluation results for a spontaneous speech corpus with about 10,000 utterances. We observed a nearly linear relationship between word accuracy and concept accuracy.
No takes yet. Share an insight, caveat, or question.
Boros et al. (2002) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: