Key points are not available for this paper at this time.
Effective valorization of construction and demolition waste (CDW) is vital for sustainable development, yet current sorting methods remain labour-intensive and inefficient. Despite progress in computer vision, large vision foundation models (LVFMs) exhibit limitations in class-specific segmentation for dynamic waste streams, while their fine-tuning demands significant computational resources. Moreover, efficient annotation of diverse waste classes remains underdeveloped, necessitating a heavy reliance on resource-intensive manual annotation for all model development. To tackle this problem, this study proposes an innovative solution by enhancing LVFMs with spatially aware adapter architectures to enable precise class-specific segmentation and introducing a semi-automated annotation pipeline leveraging few-shot learning. Experiments on an extended dataset, including underrepresented classes like rubber and lights, reveal that our approach achieves an average Intersection over Union (IoU) of 0.744 in fully supervised settings, surpassing baseline adapters by 4.5%. In few-shot scenarios, it attains an IoU of 0.770 for the unseen 'lights' class using only 20 training images. The annotation pipeline produces high-quality masks with an average IoU of 0.883, which is comparable to manual annotation performance but substantially less time-consuming. This study contributes to sustainable CDW valorization by introducing a scalable, resource-efficient framework that enhances class-specific segmentation and streamlines annotation, achieving high accuracy in both supervised and few-shot settings.
Gautam et al. (Thu,) studied this question.