Identifying noise sources in exceedance-triggered audio is essential for targeted source tracing and sustainable urban social noise governance. While accurate models require massive labeled data, the acoustic complexity, high redundancy, and imbalanced class distributions of real-world recordings incur prohibitive manual annotation costs, hindering their widespread application in IoT networks. To tackle this bottleneck, we present a label-efficient active learning framework designed to minimize annotation costs by dynamically selecting the most valuable audio samples. Specifically, rather than treating uncertainty, class balance, and diversity as separate query criteria, it encodes uncertainty and dynamic class-aware learning needs into a weighted acoustic feature space, so that diversity-based selection can be performed in a unified manner. Experiments on the UrbanSound8K benchmark and a realistic exceedance-triggered monitoring dataset demonstrate consistent label-efficiency advantages over mainstream methods. Notably, our approach reaches 98% of the fully supervised upper bound on the real-world dataset while reducing the training annotation workload by 85.0% compared to random sampling. On the real-world dataset, the proposed framework yields higher F1-scores for several challenging under-represented categories and reduces the misclassification of dominant sound events relevant to social noise source tracing. Furthermore, cross-site generalization experiments reveal rapid localized adaptation to new monitoring environments, reaching the fully supervised upper bound with only 13% of the target-domain training data. Overall, this study provides a scalable and cost-effective classification framework for urban noise monitoring, offering practical support for noise regulatory authorities and city managers in more targeted noise source tracing and governance.
Zhan et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: