PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2009199 citationsOpen Access

Fluency, adequacy, or HTER?

MSMatthew SnoverNMNitin MadnaniBDBonnie J. Dorr

Key Points

Key points are not available for this paper at this time.

Abstract

Automatic Machine Translation (MT) evaluation metrics have traditionally been evaluated by the correlation of the scores they assign to MT output with human judgments of translation performance. Different types of human judgments, such as Fluency, Adequacy, and HTER, measure varying aspects of MT performance that can be captured by automatic MT metrics. We explore these differences through the use of a new tunable MT metric: TER-Plus, which extends the Translation Edit Rate evaluation metric with tunable parameters and the incorporation of morphology, synonymy and paraphrases. TER-Plus was shown to be one of the top metrics in NIST's Metrics MATR 2008 Challenge, having the highest average rank in terms of Pearson and Spearman correlation. Optimizing TER-Plus to different types of human judgments yields significantly improved correlations and meaningful changes in the weight of different types of edits, demonstrating significant differences between the types of human judgments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Snover et al. (2009) studied this question.

synapsesocial.com/papers/6a0903cd5405cc787b9d14fdhttps://doi.org/10.3115/1626431.1626480
Ask AI
Helpful
Bookmark
Share
View Full Paper