PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 201597 citationsOpen Access

Accurate Evaluation of Segment-level Machine Translation Metrics

YGYvette GrahamTBTimothy BaldwinNMNitika Mathur

Key Points

Key points are not available for this paper at this time.

Abstract

Evaluation of segment-level machine translation metrics is currently hampered by: (1) low inter-annotator agreement levels in human assessments; (2) lack of an effective mechanism for evaluation of translations of equal quality; and (3) lack of methods of significance testing improvements over a baseline. In this paper, we provide solutions to each of these challenges and outline a new human evaluation methodology aimed specifically at assessment of segment-level metrics. We replicate the human evaluation component of WMT-13 and reveal that the current state-of-the-art performance of segment-level metrics is better than previously believed. Three segment-level metrics -METEOR, NLEPOR and SENTBLEU-MOSES -are found to correlate with human assessment at a level not significantly outperformed by any other metric in both the individual language pair assessment for Spanish-to-English and the aggregated set of 9 language pairs.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Graham et al. (2015) studied this question.

synapsesocial.com/papers/6a0e9cf506ecbe8334479d30https://doi.org/10.3115/v1/n15-1124
Ask AI
Helpful
Bookmark
Share
View Full Paper