This paper discusses the construction of a corpus for the evaluation of algorithms that generate referring expressions. It is argued that such an evaluation task requires a semantically transparent corpus, and controlled experiments are the best way to create such a resource. We address a number of issues that have arisen in an ongoing evaluation study, among which is the problem of judging the output of GRE algorithms against a human gold standard.
No takes yet. Share an insight, caveat, or question.
Deemter et al. (2006) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: