Machine translation can be evaluated using precision, recall, and the F-measure. These standard measures have significantly higher correlation with human judgments than recently proposed alternatives. More importantly, the standard measures have an intuitive interpretation, which can facilitate insights into how MT systems might be improved. The relevant software is publicly available.
No takes yet. Share an insight, caveat, or question.
Melamed et al. (2003) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: