Generative AI enhances object detection and quantification of road distress in smartphones, suggesting improved infrastructure maintenance.
Effective road distress detection and quantification are essential for ensuring transportation safety, optimizing infrastructure maintenance, and extending the service life of road networks. However, current automated systems face significant limitations due to insufficient annotated data, poor adaptability under complex environmental conditions, computational inefficiency on mobile platforms, and inaccurate physical measurement of distress features. To address these limitations, this dissertation proposes a generative AI-based framework that combines data generation, environmental adaptation, efficient mobile deployment, and three-dimensional quantification to enable integrated and field-ready road distress assessment. To mitigate data scarcity and annotation cost, a diffusion-based generative model, RoadDiffBox, was developed to synthesize diverse and class-controllable road distress images with automatic bounding-box annotations, supporting object detection training without manual labeling and achieving a mAP@50 of 0.95 with an F1-score of 0.91. Experimental results show improved dataset diversity and downstream performance, with stable generalization across geographic regions and application domains. To improve detection robustness in low-light scenarios, which often coincide with adverse environmental conditions, IllumiShiftNet was proposed. This model performs real-time nighttime-to-daytime image translation using unpaired data and a distress-focused loss function, thereby enabling accurate distress identification without requiring nighttime-specific training. It achieves high reconstruction quality, with a peak signal-to-noise ratio of 28.5 and a structural similarity index measure of 0.78, while maintaining detection performance across variable illumination levels and weather conditions. Due to the computational constraints of mobile platforms, MobiLiteNet was developed as a lightweight detection framework optimized for real-time inference on smartphones and mixed reality systems. The framework incorporates efficient channel attention, structural refinement, sparse knowledge distillation, structured pruning, and quantization to balance detection accuracy with computational efficiency, achieving processing time reductions of 71.8% on smartphones and 86.1% on mixed reality systems. For accurate dimensional quantification, MonoCrackDepthNet was developed to perform simultaneous crack segmentation and monocular depth estimation. This framework enables 3D surface reconstruction and extraction of physical dimensions from single RGB images, eliminating the need for external calibration devices or multi-sensor setups. It achieves competitive segmentation accuracy and dimensional measurement precision, with a mean relative error of 8.75% in physical quantification evaluations. In summary, this dissertation establishes an integrated, scalable, and field-deployable solution for road distress identification and quantification under real-world conditions characterized by limited data availability and environmental variability. The proposed framework advances the state of automated pavement inspection by reducing reliance on manual data collection, enhancing environmental robustness, supporting mobile deployment, and enabling accurate physical measurement, thus facilitating real-time, integrated solutions within intelligent transportation systems.
No takes yet. Share an insight, caveat, or question.
HU Yuanyuan (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: