Key points are not available for this paper at this time.
This paper investigates whether ROUGE, a popular metric for the evaluation of au-tomated written summaries, can be ap-plied to the assessment of spoken sum-maries produced by non-native speakers of English. We demonstrate that ROUGE, with its emphasis on the recall of infor-mation, is particularly suited to the as-sessment of the summarization quality of non-native speakers ’ responses. A stan-dard baseline implementation of ROUGE-1 computed over the output of the au-tomated speech recognizer has a Spear-man correlation of ρ = 0.55 with experts’ scores of speakers ’ proficiency (ρ = 0.51 for a content-vector baseline). Further in-creases in agreement with experts ’ scores can be achieved by using types instead of tokens for the computation of word fre-quencies for both candidate and reference summaries, as well as by using multiple reference summaries instead of a single one. These modifications increase the cor-relation with experts ’ scores to a Spear-man correlation of ρ = 0.65. Furthermore, we found that the choice of reference sum-maries does not have any impact on per-formance, and that the adjusted metric is also robust to errors introduced by auto-mated speech recognition (ρ = 0.67 for hu-man transcriptions vs. ρ = 0.65 for speech recognition output). 1
Loukina et al. (Wed,) studied this question.