PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 25, 2026Journal of Data and Information Quality0 citationsOpen Access

Improving Document Quality via Updatable Extracted Views

View Full Paper
BKBesat KassaieUniversity of WaterlooFTFrank TompaUniversity of Waterloo

Key Points

  • This research aims to enhance data quality in unstructured documents using extracted views for targeted data cleaning.
  • Utilized rule-based extraction programs to generate tabular views from unstructured documents.
  • Characterized sufficient conditions that allow document updates without affecting the quality of the extracted views.
  • Conducted experiments on medical records and math-intensive documents to validate the proposed approach.
  • Document updates under the proposed system do not introduce unintended changes to extracted views, ensuring document cleanliness.
  • The math retrieval system exhibited a performance improvement of at least 100% on cleaned documents compared to uncleaned ones, matching tailored cleaning methods.
  • The approach facilitates the preservation of cleaned documents for further applications.

Abstract

Data cleaning is a critical component of modern data intelligence systems. However, most existing approaches have been proposed for tabular data, despite the fact that much real-world data is unstructured. We propose to leverage information extraction programs to disclose where to apply data cleaning processes for documents. After using information extraction to produce a table of values extracted from the document (an “extracted view” of the documents, in the sense of a database view ), the resulting table is subject to any desired tabular-based data cleaning protocol. However, in order to reflect such cleaning back in the source documents, we require that updates to the tabular views can be translated into document updates without introducing any unintended changes to the contents in those views. With this approach, we consider a document to be clean if all its extracted views are clean. In this paper, we characterize and verify a set of sufficient conditions for rule-based extraction programs that ensure document updates do not introduce unintended changes to the extracted views, thus qualifying them for inclusion in a document cleaning pipeline. Through experiments conducted on medical records and on math-intensive documents, we demonstrate that our approach provides an effective, practical pipeline for correcting data quality problems in documents. The second of these experiments demonstrates that a math retrieval system can perform at least twice as well on the cleaned documents as it does when applied to the documents without cleaning, matching the improvements achieved by specially tailored cleaning methods applied to those documents during the search pipeline but preserving the cleaned documents for use in other applications as well.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kassaie et al. (2026) studied this question.

synapsesocial.com/papers/6a13e8030e02ee3982d32b38https://doi.org/10.1145/3815119
Ask AI
Helpful
Bookmark
Share
View Full Paper