PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 20250 citationsOpen Access

Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos

View Full Paper
DBDavide BerghiPJPhilip J. B. Jackson

Key Points

  • The method significantly improves sound event localization in videos by incorporating spatial and semantic embeddings.
  • Results showed a notable performance boost, achieving second rank in the DCASE 2025 Challenge Task 3.
  • The study introduces the Cross-Modal Conformer, enhancing traditional architectures with multimodal fusion.
  • Synthetic datasets were curated and expanded, underpinning the need for large-scale training in 3D SELD tasks.

Abstract

In this study, we address the multimodal task of stereo sound event localization and detection with source distance estimation (3D SELD) in regular video content. 3D SELD is a complex task that combines temporal event classification with spatial localization, requiring reasoning across spatial, temporal, and semantic dimensions. The last is arguably the most challenging to model. Traditional SELD approaches typically rely on multichannel input, limiting their capacity to benefit from large-scale pre-training due to data constraints. To overcome this, we enhance a standard SELD architecture with semantic information by integrating pre-trained, contrastive language-aligned models: CLAP for audio and OWL-ViT for visual inputs. These embeddings are incorporated into a modified Conformer module tailored for multimodal fusion, which we refer to as the Cross-Modal Conformer. We perform an ablation study on the development set of the DCASE2025 Task3 Stereo SELD Dataset to assess the individual contributions of the language-aligned models and benchmark against the DCASE Task 3 baseline systems. Additionally, we detail the curation process of large synthetic audio and audio-visual datasets used for model pre-training. These datasets were further expanded through left-right channel swapping augmentation. Our approach, combining extensive pre-training, model ensembling, and visual post-processing, achieved second rank in the DCASE 2025 Challenge Task 3 (Track B), underscoring the effectiveness of our method. Future work will explore the modality-specific contributions and architectural refinements.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Berghi et al. (2025) studied this question.

synapsesocial.com/papers/68ec1be02b8fa9b2b78ad11chttps://doi.org/10.48550/arxiv.2509.06598
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection2024 · 14 citations
  2. 2Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification2025
  3. 3A Gated Multi-Head Network with Semantic–Spatial Features for Sound Event Localization and Detection2026
  4. 4Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios2024
  5. 5Sound Event Detection and Localization with Distance Estimation2024