Sentence level evaluation in MT has turned out far more difficult than corpus level evaluation. Existing sentence level metrics employ a lim-ited set of features, most of which are rather sparse at the sentence level, and their intricate models are rarely trained for ranking. This pa-per presents a simple linear model exploiting 33 relatively dense features, some of which are novel while others are known but seldom used, and train it under the learning-to-rank frame-work. We evaluate our metric on the stan-dard WMT12 data showing that it outperforms the strong baseline METEOR. We also ana-lyze the contribution of individual features and the choice of training data, language-pair vs. target-language data, providing new insights into this task. 1
No takes yet. Share an insight, caveat, or question.
Stanojević et al. (2014) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: