Key points are not available for this paper at this time.
Vision-based monitoring of personal protective equipment (PPE) is central to construction safety, yet robust detectors remain limited by scarce, privacy-constrained site imagery. Digital twin simulation can generate labeled synthetic data at scale, but Sim-to-Real gaps make the effective use of synthetic data under a fixed training budget unclear. We benchmark YOLOv11s, Faster R-CNN, and RT-DETR-L using a controlled real–synthetic mixing protocol comprising a fixed real-only test set (400 images), separate sampling pools (5760 real and 4000 synthetic images), and eleven training configurations of approximately constant size (∼1450 images before the validation split) with real fractions ranging from 0% to 100%. Using average recall (AR@100) as the primary safety-oriented metric, the original single-run benchmark shows non-linear architecture-dependent responses to data mixing: YOLOv11s and RT-DETR-L achieve their single-run peaks at G9 (90% real/10% synthetic), whereas Faster R-CNN performs best at G10 (100% real). To assess robustness for the most central YOLOv11s comparison, we further conduct a targeted supplementary repeated-seed analysis for G9 and G10 and re-evaluate all resulting checkpoints on the same fixed real-only test set. This supplementary analysis shows that G10 achieves higher mean performance and lower variance than G9 for YOLOv11s, indicating that the apparent single-run advantage of limited synthetic supplementation is not stable across reruns. However, this robustness check is limited to the central YOLOv11s G9-versus-G10 case and should not be interpreted as a comprehensive robustness validation across all configurations and detector families. Persistent errors on safety vests further indicate a materiality gap for deformable PPE. Overall, these findings suggest that synthetic supplementation can be useful in some settings, but its value is architecture-dependent, evaluation setting-sensitive, and should be interpreted cautiously under robustness-oriented evaluation.
Zhang et al. (Thu,) studied this question.