PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 8, 2026Information Systems2 citationsOpen Access

Process mining between the lines: Extracting object-centric event logs from textual data

View Full Paper
ABAlina BussTechnical University of MunichCKChristoph KechtFraunhofer Institute for Applied Information TechnologyWKWolfgang KratschTechnische Hochschule Augsburg

Key Points

  • This research aims to develop a method for extracting object-centric event logs from unstructured textual data to enhance process mining.
  • Developed an automated extraction approach comprising a collector and a refiner.
  • Utilized heuristic natural language processing and generative artificial intelligence techniques.
  • Architectured multiple combinations of collector and refiner instances for evaluation.
  • Conducted assessments on both synthetic texts and natural corpora, including fire status updates and a legal judgment.
  • The generative collector configurations achieved the highest extraction quality.
  • The fully generative variant produced coherent and standardized event and object labels.
  • Demonstrated practical utility on two distinct corpora for process mining applications.

Abstract

Organizations generate vast amounts of unstructured textual data – a valuable source of information that frequently remains underutilized for process mining. However, textual descriptions often record exceptions and manual activities absent from structured data, and therefore, enable a better understanding of deviations from the expected business process behavior. Importantly, unstructured sources typically retain the object-centric characteristics of real-world processes – information that gets flattened or lost in case-centric event logs. Yet, existing approaches primarily target structured data sources or produce case-centric event logs. To address this gap, we present an automated approach to derive object-centric event logs directly from unstructured textual descriptions. The approach comprises two subcomponents: a collector that identifies events and objects (including their attributes and relationships), and a refiner that consolidates and cleans the extracted information. We instantiate each subcomponent in heuristic and generative implementations and create four pairwise combinations of collector and refiner instances to assess the effectiveness of heuristic natural language processing and generative artificial intelligence techniques. We compare these variants quantitatively and qualitatively in a controlled, artificial setting based on synthesized texts and demonstrate the practical utility on two naturally occurring corpora (fire status updates and a legal judgment). Our results show that the configurations with a generative collector achieve the highest extraction quality. In particular, the fully generative variant produces coherent and standardized event and object labels. Overall, this study fills a notable research gap by enabling the incorporation of textual information into process mining applications. • Proposes an approach to extract object-centric event logs from textual descriptions. • Develops the approach using the Design Science Research methodology. • Implements the approach using heuristic NLP and generative AI techniques. • Evaluates the approach on synthetic and naturalistic textual descriptions. • Confirms the approach’s practical utility on fire status updates and a legal judgment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Buss et al. (2026) studied this question.

synapsesocial.com/papers/69ada962bc08abd80d5bc981https://doi.org/10.1016/j.is.2026.102713
Ask AI
Helpful
Bookmark
Share
View Full Paper