PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 29, 2026Scientific Data1 citationsOpen Access

Multimodal and Hyperspectral Dataset for Segmentation of Bulky Waste using VIS, IR, NIR, and Terahertz Imaging

MBManuel BihlerKarlsruhe Institute of TechnologyLRLukas RomingFraunhofer Institute of Optronics, System Technologies and Image ExploitationDČDovilė ČibiraitėFraunhofer Institute for Industrial Mathematics

Key Points

  • To provide a comprehensive dataset for segmentation and classification of bulky waste using multiple imaging modalities.
  • Annotated dataset created from four imaging modalities: VIS, NIR, IR, and THz.
  • Image registration aligns different modalities for enhanced accuracy.
  • Dataset includes 22,659 annotated patches for wood and non-wood classes.
  • Predefined splits for training, validation, and testing were established.
  • Baseline performance reported using convolutional neural networks.
  • Dataset contains 56 multi-sensor scenes for robust evaluations.
  • Binary discrimination task revealed effective segmentation of wood versus non-wood.
  • Challenges identified include occlusion and concealed contaminants, promoting advanced fusion techniques.

Abstract

Abstract This study presents an annotated multi-sensor, multimodal, and hyperspectral dataset designed to support deep learning-based classification and segmentation of bulky waste. The dataset comprises four distinct sensor modalities: high-resolution visible RGB images (VIS), hyperspectral near-infrared (NIR), temporally resolved thermal infrared (IR), and terahertz (THz) imaging with depth information, providing complementary multimodal information. An image registration process aligns all modalities to a common reference frame, enabling near pixel-precise fusion across sensors. WoodVIT contains 56 registered multi-sensor scenes, partitioned into 22,659 annotated patches with two main classes (wood and non-wood) and 16 subclass labels. It includes pixel-masks and patch-wise annotations to facilitate both segmentation and classification tasks. The primary benchmark task is binary discrimination of wood versus non-wood. The dataset also includes challenging scenarios involving occlusion and concealed contaminants (e.g., embedded metals) to motivate robust multimodal fusion approaches. We provide predefined train/validation/test splits and report baseline results using convolutional neural networks and fusion architectures to establish reference performance. WoodVIT is publicly available to support research on multi-sensor learning for waste sorting.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bihler et al. (2026) studied this question.

synapsesocial.com/papers/69c8c2fcde0f0f753b39d8c4https://doi.org/10.1038/s41597-026-07053-1
Ask AI
Helpful
Bookmark
Share
View Full Paper