AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (989.2 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Sound event localization and detection network with enhanced feature expression

Dongping ZHANG1( )Zhentao FU1Zhutao WANG1Lili LIN2Ming WEI3
College of Information Engineering,China Jiliang University,Hangzhou 310018,China
School of Information and Electronic Engineering,Zhejiang Gongshang University,Hangzhou 310018,China
Hangzhou Aihua Intelligent Technology Co.,Ltd.,Hangzhou 311121,China
Show Author Information

Abstract

To address the problem that traditional deep learning models are difficult to capture the long-context feature correlations in input feature maps as well as the key feature information in channel and spatial dimensions, resulting in high error rates and unsatisfactory performance in sound event localization and detection (SELD). Based on the baseline model SELDnet in the acoustic scene classification and sound event detection challenge, this paper proposes a feature enhanced sound event localization and detection network (FE-SELDnet). In order to address the issue of function failure to backpropagate, which leads to neuron death, it suggests using group normalization and the SiLU activation function; introducing the convolutional block attention module (CBAM) to capture significant features in both channel and spatial dimensions of acoustic features, suppressing superfluous features, improving network sensitivity and accuracy to feature information, and improving information flow; introducing the Transformer module to capture longer speech context feature association and combine local features to improve the accuracy and robustness of the model in sound event detection and localization tasks. The proposed FE-SELDnet significantly outperforms the original baseline network, according to experimental results on the TUT Sound Events dataset. The error rate decreased from 0.45 to 0.326, the SED and DOA scores decreased from 0.45 and 0.32 to 0.26 and 0.25, respectively, and the F1 score increased to 79.4%. The algorithm proposed in this paper has higher superiority.

CLC number: TN912.3;TP181 Document code: A Article ID: 1001-5965(2026)04-1088-08

References

【1】
【1】
 
 
Journal of Beijing University of Aeronautics and Astronautics
Pages 1088-1095

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
ZHANG D, FU Z, WANG Z, et al. Sound event localization and detection network with enhanced feature expression. Journal of Beijing University of Aeronautics and Astronautics, 2026, 52(4): 1088-1095. https://doi.org/10.13700/j.bh.1001-5965.2024.0019

192

Views

0

Downloads

0

Crossref

1

Scopus

0

CSCD

Received: 11 January 2024
Published: 15 March 2024
© Journal of Beijing University of Aeronautics and Astronautics