Following the recent adoption by the machine translation community of automatic evaluation using the BLEU/NIST scoring process, we conduct an in-depth study of a similar idea for evaluating summaries. The results show that automatic evaluation using unigram co-occurrences between summary pairs correlates surprising well with human evaluations, based on various statistical metrics; while direct application of the BLEU evaluation procedure does not always give good results.
No takes yet. Share an insight, caveat, or question.
Lin et al. (2003) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: