Dialogue summarization remains a challenging task due to conversational ambiguity and informality. This work explores the relationship between human evaluation scores and automatic metrics for thirty dialogue summaries, and aims to identify whether these metrics represent human quality perception. Human evaluators rated summaries on accuracy, conciseness, meaning preservation, and grammar using a 5-point scale. Scores were compared with ROUGE-1, ROUGE-2, ROUGE-L, and BERTScore. Results show a moderate alignment between human perception and BERTScore, while ROUGE metrics show weaker correlations. These findings suggest reference metrics capture semantic alignment more effectively than structural quality. Recommendations for improving summarization include incorporating factual consistency and compression-aware training. The study highlights the continued importance of human evaluation in conversational summarization research.
Yunus et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: