PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 17, 20240 citationsOpen Access

Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4

View Full Paper
SSSangwon SonJPJ.B. ParkHKHong-Kook Kim

Key Points

Key points are not available for this paper at this time.

Abstract

In this report, we propose three novel methods for developing a sound event detection (SED) model for the DCASE 2024 Challenge Task 4. First, we propose an auxiliary decoder attached to the final convolutional block to improve feature extraction capabilities while reducing dependency on embeddings from pre-trained large models. The proposed auxiliary decoder operates independently from the main decoder, enhancing performance of the convolutional block during the initial training stages by assigning a different weight strategy between main and auxiliary decoder losses. Next, to address the time interval issue between the DESED and MAESTRO datasets, we propose maximum probability aggregation (MPA) during the training step. The proposed MPA method enables the model's output to be aligned with soft labels of 1 s in the MAESTRO dataset. Finally, we propose a multi-channel input feature that employs various versions of logmel and MFCC features to generate time-frequency pattern. The experimental results demonstrate the efficacy of these proposed methods in a view of improving SED performance by achieving a balanced enhancement across different datasets and label types. Ultimately, this approach presents a significant step forward in developing more robust and flexible SED models

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Son et al. (2024) studied this question.

synapsesocial.com/papers/68e64779b6db6435875d931chttps://doi.org/10.48550/arxiv.2406.12721
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1FMSG-JLESS Submission for DCASE 2024 Task4 on Sound Event Detection with Heterogeneous Training Dataset and Potentially Missing Labels2024
  2. 2Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries2025
  3. 3MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection2024 · 1 citations
  4. 4DiffSED: Sound Event Detection with Denoising Diffusion2024 · 10 citations
  5. 5Gate‐Align‐SED: Semi‐Supervised Sound Event Detection via Adaptive Feature Gating and Cross‐Task Alignment in Situation Awareness2026