PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 8, 2026Water Resources Research3 citationsOpen Access

Exploring the Transferability of Image‐Based Algorithms for River Plastic Detection: The Value of Small Mixed Data Sets

View Full Paper
KSKhim Cathleen SaddiTETim van EmmerikDMDomenico Miglino

Key Points

  • This research aims to evaluate the efficiency of small, diverse training sets for detecting river plastic using image-based algorithms.
  • Compiled four river-camera data sets from multiple countries.
  • Trained YOLOv7 and YOLOv8 models to assess detection performance.
  • Evaluated the impact of data diversity and class aggregation on detection outcomes.
  • Employed internal and external validation to examine transferability.
  • Diverse data sets showed higher performance per annotation compared to homogeneous data sets.
  • Naïve data merging can degrade internal validation results without careful filtering.
  • Class aggregation improved overall detection skill significantly.
  • Internal validation alone is not a reliable predictor of performance at external sites.

Abstract

Abstract Using large image data sets has been the conventional strategy to improve object detection, but it increases annotation effort and training cost and does not guarantee robust transfer to new sites. Here we quantify the value of a small, diverse training set for floating macroplastic detection by jointly evaluating performance, computational cost, annotation effort, and cross‐site transferability. We compile four river‐camera data sets from Indonesia, The Netherlands, and Vietnam (training/internal validation) and Italy (external validation, single site and single day data from a long‐term camera monitoring system), harmonized into 13 litter classes and a five‐level tiering scheme (progressive class aggregation). We train YOLOv7 and YOLOv8 models and compare site‐specific data sets with a merged “Mixed” data set (999 images) spanning heterogeneous environmental conditions. Results show that data sets with more diverse backgrounds (Type II; e.g., The Netherlands) achieve higher performance per annotation than homogeneous data sets (Type I; e.g., Indonesia, Vietnam), whereas naïvely merging data sets can degrade internal validation unless accompanied by feature‐aware filtering. Class aggregation substantially increases overall detection skill, with gains consistent across data sets when moving from fine (Tier 4) to coarse (Tier 0) label spaces. Finally, internal validation does not reliably predict external‐site performance, underscoring the need for transferability‐aware data set design and evaluation. Overall, our findings emphasize that data diversity and curation , rather than data set size alone, are key levers to scale river plastic detection toward broader deployment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Saddi et al. (2026) studied this question.

synapsesocial.com/papers/69d5f00974eaea4b11a798f1https://doi.org/10.1029/2025wr040605
Ask AI
Helpful
Bookmark
Share
View Full Paper