This study investigates the improvement in recognizing and categorizing acoustic events in complex audio environments through a combination of convolutional neural networks (CNN), log-mel spectrograms, and long-short-term memory (LSTM) networks. By leveraging these algorithms, contextual information and temporal dependencies are captured, facilitating the extraction of time-frequency features from audio data. Various deep learning models were evaluated, with the proposed CNN+LSTM architecture emerging as the top performer. Achieving an accuracy of 91% and 88% for one-fold validation and 10-fold cross-validation, respectively, on the UrbanSound8k dataset, underscores the efficacy of our approach compared to recent techniques in sound-event detection. Notably, the CNN+LSTM model offers a lightweight alternative to RNNs and pre-trained models, holding considerable promise for applications in environmental monitoring and audio surveillance domains.
No takes yet. Share an insight, caveat, or question.
Shanmukha et al. (2024) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: