PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 20250 citationsOpen Access

LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis

View Full Paper
IHIn-Wook HeoTHTaewook HwangJJJeesu Jung

Key Points

  • The LED benchmark effectively identifies critical structural errors in document layouts, enhancing evaluation accuracy.
  • Results indicate that LED differentiates layout understanding capabilities across various large language models.
  • Traditional evaluation metrics fall short in detecting errors like region merging and content omission.
  • The synthetic LED-dataset was created to simulate realistic structural errors based on empirical distributions.

Abstract

Recent advancements in Document Layout Analysis through Large Language Models and Multimodal Models have significantly improved layout detection. However, despite these improvements, challenges remain in addressing critical structural errors, such as region merging, splitting, and missing content. Conventional evaluation metrics like IoU and mAP, which focus primarily on spatial overlap, are insufficient for detecting these errors. To address this limitation, we propose Layout Error Detection (LED), a novel benchmark designed to evaluate the structural robustness of document layout predictions. LED defines eight standardized error types, and formulates three complementary tasks: error existence detection, error type classification, and element-wise error type classification. Furthermore, we construct LED-Dataset, a synthetic dataset generated by injecting realistic structural errors based on empirical distributions from DLA models. Experimental results across a range of LMMs reveal that LED effectively differentiates structural understanding capabilities, exposing modality biases and performance trade-offs not visible through traditional metrics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Heo et al. (2025) studied this question.

synapsesocial.com/papers/68e70db790569dd607ee677dhttps://doi.org/10.48550/arxiv.2507.23295
Ask AI
Helpful
Bookmark
Share
View Full Paper