Construction sites implement various safety management activities, including toolbox meetings, risk assessments, and safety knowledge assessments, to reduce accidents. Multiple-choice question (MCQ)-based assessments are widely used to evaluate worker safety competencies. However, the effectiveness of MCQ assessments depends critically on distractor quality; incorrect options must be plausible enough to challenge uninformed respondents while remaining clearly distinguishable from knowledgeable ones. Manual distractor creation requires substantial expertise and is prone to inconsistency, whereas large language models (LLMs) often generate options that lack domain relevance. This paper proposes context-aware multipath adaptive safety scoring (CoMPASS), an algorithm that integrates construction safety domain knowledge with LLM capabilities for MCQ distractor generation. CoMPASS operates through two pathways: CoMPASS-H leverages a hierarchical hazard factor ontology for hazard identification questions, whereas CoMPASS-R uses hybrid retrieval-augmented generation (RAG) for risk control questions. An evaluation using 50 real construction accident cases with a robotic assessment test (RAT) using frontier LLMs as virtual examinees demonstrated that CoMPASS-R achieved a 90% quality pass rate, whereas all baseline methods failed to meet the composite quality criteria. The proposed framework provides a scalable approach to generating assessment content that supports effective safety management at construction sites.
Shin et al. (Thu,) studied this question.