Randomized trial explores correlation discovery between structured and unstructured data in automotive data lakes, suggesting improved data management techniques.
Data Lakes have emerged as an architectural approach for integrating data from multiple heterogeneous sources into large-scale data processing pipelines. A key challenge lies in correlating structured, semi-structured, and unstructured data through metadata in order to transform raw data into actionable information and support automated decision-making at scale. This research paper builds upon SmartData, a self-contained data construct originally designed for structured data, to propose a Data Lake infrastructure that enables efficient integration across diverse data modalities. We introduce SmartDataContext, an intermediary semi-structured representation that encapsulates SmartData Models and supports time-series tagging. Additionally, we propose a SmartTagging framework that extracts semantic information from unstructured data and applies text relationship analysis to automatically discover correlations. The proposed approach is evaluated through an automotive case study, demonstrating the identification of relationships between sections of standards documents in PDF format and corresponding SmartDataContexts. From a data management viewpoint, SmartTagging acts as an “active metadata” mechanism for automotive data lakes, allowing for automatic and continuous discovery of tags from contextual artifacts and external standards (here, ETSI C-ITS documents), thus, avoiding manual catalogs or schema-on-read and the creation of data-swamps.
No takes yet. Share an insight, caveat, or question.
Gonçalves et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: