Machine learning study demonstrates enhanced accident information retrieval using graph retrieval-augmented generation in railway safety profiles, indicating improved safety management systems.
In railway operations and safety oversight, vast amounts of accident-related data are recorded in unstructured textual formats, posing challenges for efficient information extraction and analysis. To address this, we constructed a knowledge graph from railway accident profile texts and integrated it with a Graph Retrieval-Augmented Generation (GraphRAG)-enhanced retrieval framework to support both basic and composite queries on railway accident information. First, railway accident profile texts were preprocessed, and Easy Data Augmentation (EDA) was used to expand samples of different types, thereby constructing a railway accident profile text dataset for subsequent knowledge extraction. We then introduced a deep learning model, RoBERTa-CNN-BiLSTM-CRF (RCBC), for automatic extraction of seven categories of entities, including accident IDs and causes. To extract semantic relations, we designed a prompt-based template leveraging large language models (LLMs). To mitigate information loss during entity extraction, a GCN-Attention-LLM (GA-LLM) model was further designed for knowledge graph completion. Experimental results show that RCBC achieves MicroF scores above 85% across entity extraction tasks, while GA-LLM attains an average Hits@3 of 82.84% in knowledge completion. In tests on basic and composite questions, LLM-GraphRAG outperformed both LLM and LLM-RAG in faithfulness, semantic similarity, context precision, and context recall. The resulting knowledge graph contains 1493 entities and 1832 relations. Combined with the retrieval framework, the system enables access to key railway accident information and offers technical support for intelligent railway safety management.
No takes yet. Share an insight, caveat, or question.
Lian et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: