evaluation measures produce a ranking of all possible extract summaries of a document., Recall-based evaluation measures, which depend on costly human-generated ground truth summaries, produce uncorrelated rankings when ground truth is varied. This paper proposes using sentence-rankbased and content-based measures for evaluating extract summaries, and compares these with recallbased evaluation measures. Content-based measures increase the correlation of rankings induced by synonymous ground truths, and exhibit other desirable properties.
No takes yet. Share an insight, caveat, or question.
Donaway et al. (2000) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: