Automatically constructed datasets for generating text from semi-structured data (tables), such as WikiBio We show that metrics which rely solely on the reference texts, such as BLEU and ROUGE, show poor correlation with human judgments when those references diverge. We propose a new metric, PAR-ENT, which aligns n-grams from the reference and generated texts to the semi-structured data before computing their precision and recall. Through a large scale human evaluation study of table-to-text models for WikiBio, we show that PARENT correlates with human judgments better than existing text generation metrics. We also adapt and evaluate the information extraction based evaluation proposed in We show that PARENT is also applicable when the reference texts are elicited from humans using the data from the WebNLG challenge.
No takes yet. Share an insight, caveat, or question.
Dhingra et al. (2019) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: