This paper presents a methodology for the evaluation of table understanding algorithms for PDF documents. The evaluation takes into account three major tasks: table detection, table structure recognition and functional analysis. We provide a general and flexible output model for each task along with corresponding evaluation metrics and methods. We also present a methodology for collecting and ground-truthing PDF documents based on consensus-reaching principles and provide a publicly available ground-truthed dataset.
No takes yet. Share an insight, caveat, or question.
Göbel et al. (2012) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: