PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 25, 2026Buildings0 citationsOpen Access

Time-Aware Construction Site Risk Prediction Based on Sentence-BERT and 7-Day Window Aggregation with Unlabeled Data

View Full Paper
SLShuo LiuWYWeidong YanGLGuoqi Liu

Key Points

  • The research aims to enhance risk identification and prioritization in construction through semantic analysis of safety texts.
  • Constructed a keyword library and extracted risk terms using tokenization and TF-IDF.
  • Utilized a fine-tuned Sentence-BERT model for generating sentence embeddings.
  • Applied FAISS-based similarity search to link inspection records with historical accidents.
  • Introduced a 7-day inspection window to assess hazard accumulation.
  • Achieved an accuracy of 0.75 and a recall of 1.00.
  • F1-score of 0.8571 in the primary tests.
  • Cross-project validation yielded an F1-score of 0.5607.
  • The framework remained robust against 10% noise interference.

Abstract

Construction safety texts are commonly used only for descriptive statistical analysis, and systematic approaches for uncovering latent semantic risk correlations remain limited. In particular, risk identification and prioritization under unlabeled conditions remain challenging. To address this issue, this study proposes a semantic risk association and ranking framework based on Sentence-BERT (SBERT). First, a domain-specific keyword library is constructed, and representative risk terms are extracted through tokenization, stop-word removal, and TF-IDF weighting. A fine-tuned SBERT model is then employed to generate sentence embeddings. FAISS-based similarity search is applied to match safety inspection records with historical accident reports, enabling automatic identification and ranking of the most relevant accident types. In addition, a seven-day inspection window is introduced to capture the temporal accumulation effect of hazards and support risk assessment without explicit labels. Experiments conducted on 1368 accident reports and 484 inspection records show that the proposed framework achieves an accuracy of 0.75, a recall of 1.00, and an F1-score of 0.8571. Cross-project validation yields an F1-score of 0.5607, and the performance remains stable under 10% noise interference. The results demonstrate that the proposed semantic risk association and ranking framework is effective and robust for practical construction safety management.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2026) studied this question.

synapsesocial.com/papers/69c37be2b34aaaeb1a67eae4https://doi.org/10.3390/buildings16061243
Ask AI
Helpful
Bookmark
Share
View Full Paper