PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026IEEE Access2 citationsOpen Access

Environmental Sound Classification Using Advanced Pooling and Hybrid Convolution-KAN Architecture

View Full Paper
PSPasan SarathchandraDMDimalsha MadushaniSMSenuri Mallikarachchi

Key Points

  • Accuracy of 88.25% on ESC-50 confirms the model's robustness in environmental sound classification tasks.
  • The methodology integrates convolutional neural networks with Kolmogorov-Arnold Networks for enriched feature representation.
  • Dimensionality reduction is achieved through advanced pooling techniques including PCA and Sparse Salient Region Pooling.
  • Potential scalability of this model design could lead to broader applications in smart cities and IoT systems.

Abstract

Environmental Sound Classification plays a crucial role in diverse applications ranging from smart cities and forest monitoring to surveillance and context-aware IoT systems. Convolutional Neural Networks have recently emerged as the dominant paradigm, surpassing traditional approaches in environmental sound classification tasks. However, these performance gains often come at the cost of increased network depth, model complexity, and model size, limiting their usage in many practical applications. To address these challenges, we present a novel hybrid deep learning architecture that combines CNNs, Kolmogorov-Arnold Networks (KAN), and advanced pooling strategies such as Sparse Salient Region Pooling (SSRP) and Principal Component Analysis (PCA) pooling. Our methodology follows a progressive enhancement strategy: starting with a baseline CNN model, we integrate KAN layers for richer functional representation, introduce SSRP to better capture salient regions, and replace conventional pooling with PCA pooling for dimensionality reduction. To ensure robust feature learning, we adopt a multi-stage preprocessing pipeline consisting of waveform-level augmentations (pitch shifting, time stretching, Gaussian noise addition, random gain), log-mel spectrogram feature extraction, and feature-level augmentation techniques including time and frequency masking and mixup. Experimental results demonstrate that our proposed lightweight model with approximately 500K parameters yields accuracies of 88.25% on ESC-50 and 96.75% on ESC-10, 85.13% on UrbanSound8K and 87.92% on FSC22 datasets. This establishes a strong balance between performance and model compactness, and outperforming existing state-of-the-art approaches. The results highlight the potential of hybrid neural designs that fuse convolutional front-ends, operator-based back-ends, and adaptive pooling mechanisms for environmental sound classification.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sarathchandra et al. (2026) studied this question.

synapsesocial.com/papers/69a760a2c6e9836116a2d8ffhttps://doi.org/10.1109/access.2026.3660484
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Deep Learning2023 · 36 citations
  2. 2Convolutional Neural Networks: A Comprehensive Evaluation and Benchmarking of Pooling Layer Variants2024 · 18 citations
  3. 3An efficient medical image classification network based on multi-branch CNN, token grouping Transformer and mixer MLP2024 · 74 citations
  4. 4Acoustic Event Classification with Enhanced EfficientNet2024 · 4 citations
  5. 5Environmental Sound Classification With Low-Complexity Convolutional Neural Network Empowered by Sparse Salient Region Pooling2022 · 19 citations