PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 1, 2026Semantic Web2 citationsOpen Access

Incremental Knowledge Graph Construction from Heterogeneous Data Sources

View Full Paper
DADylan Van AsscheJRJulián Andrés RojasBMBen De Meester

Key Points

  • This research aims to improve the efficiency of knowledge graph (KG) construction from datasets that frequently change.
  • Investigated change signaling strategies from real-world datasets.
  • Proposed change detection algorithms tailored for different data sources.
  • Introduced a declarative approach using RDF Mapping Language (RML) to apply changes to KGs.
  • Implemented the approach in RMLMapper as Incremental RML and evaluated it functionally and quantitatively.
  • Achieved up to 315.83x reduction in storage for multiple KG versions.
  • Reduced CPU time by 4.59x and memory usage by 1.51x during KG generation.
  • Decreased KG construction time by up to 4.41x, especially for larger datasets.

Abstract

Sharing datasets that change (through creates, updates, deletes) poses challenges to data consumers, including reconciling historical versioning and managing frequent changes. This is evident for Knowledge Graphs (KGs), materialized from such datasets, where synchronization happens through frequent regeneration. However, this is time-consuming, loses history, and wastes computing resources through redundant processing. We present a KG generation approach that efficiently handles evolving data sources with different change signaling strategies. We investigate change signaling strategies of real-world datasets, propose corresponding change detection algorithms, and introduce a declarative approach based on the RDF Mapping Language (RML) and Function Ontology to materialize changes for evolving KGs. Detected changes can be automatically published as a Linked Data Event Stream (LDES), using the Activity Streams 2.0 vocabulary to describe changes and communicate them over the Web. We implement our approach in the RMLMapper as Incremental RML and evaluate it both functionally, and quantitatively using a modified version of the GTFS Madrid Benchmark and several real-world data sources. Our approach reduces storage and computing requirements for generating and storing multiple KG versions (up to 315.83x less storage, 4.59x less CPU time, and 1.51x less memory) and reduces KG construction time up to 4.41x. Performance gains are more pronounced for larger datasets, while our approach's overhead partially offsets benefits for smaller ones. Overall, our approach lowers the cost of publishing and maintaining KGs and, via LDES, supports timely, Web-native dissemination of changes. We plan to optimize our change detection algorithms and use windowing to support streaming data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Assche et al. (2026) studied this question.

synapsesocial.com/papers/69a3d7eeec16d51705d2e565https://doi.org/10.1177/22104968251412270
Ask AI
Helpful
Bookmark
Share
View Full Paper