To address the problem that traditional deep learning models are difficult to capture the long-context feature correlations in input feature maps as well as the key feature information in channel and spatial dimensions, resulting in high error rates and unsatisfactory performance in sound event localization and detection (SELD). Based on the baseline model SELDnet in the acoustic scene classification and sound event detection challenge, this paper proposes a feature enhanced sound event localization and detection network (FE-SELDnet). In order to address the issue of function failure to backpropagate, which leads to neuron death, it suggests using group normalization and the SiLU activation function; introducing the convolutional block attention module (CBAM) to capture significant features in both channel and spatial dimensions of acoustic features, suppressing superfluous features, improving network sensitivity and accuracy to feature information, and improving information flow; introducing the Transformer module to capture longer speech context feature association and combine local features to improve the accuracy and robustness of the model in sound event detection and localization tasks. The proposed FE-SELDnet significantly outperforms the original baseline network, according to experimental results on the TUT Sound Events dataset. The error rate decreased from 0.45 to 0.326, the SED and DOA scores decreased from 0.45 and 0.32 to 0.26 and 0.25, respectively, and the F1 score increased to 79.4%. The algorithm proposed in this paper has higher superiority.
- Article type
- Year
To address the limitations of existing abnormal behavior detection models—particularly their inadequate feature representation and insufficient modeling of dynamic temporal characteristics—this paper proposes a multi-channel coupled spatio-temporal enhanced anomaly detection method. Built upon the SlowFast network architecture, the proposed approach integrates a multi-channel coupled spatial enhancement module into the slow pathway to strengthen static feature modeling, and a multi-channel coupled temporal enhancement module into the fast pathway to improve the discriminability of dynamic temporal features. Extensive experiments on three benchmark datasets—Violent Flow, Hockey Fight, and Real-life Violence Situations—demonstrate that the proposed method achieves prediction accuracies of 95.3%, 97.3%, and 94%, respectively, outperforming current state-of-the-art approaches. The results validate the superior feature representation capability and generalization performance of the proposed method in abnormal behavior recognition tasks.
In order to solve the problem that traditional aircraft skin defect detection relies on human eye observation, which leads to reduced efficiency due to easy fatigue of the human eye and limited individual cognition, an aircraft skin defect detection algorithm based on improved YOLOv8 is proposed. Improve the data improvement strategy and propose a new one that combines slice reasoning with mosaic. Integrate the residual block into the feature extraction network to enhance the network expression ability and improve the accuracy of the model in aircraft skin defect detection tasks. Use the triplet attention module to strengthen the feature fusion network and lower the false and missed detection rates of small target samples. Optimize the structure of the detection head so that the network can better effectively combine shallow information with depth information. On the aircraft skin defect data set, experimental results indicate that the revised algorithm’s mean average precision (mAP) and recall rate have increased by 3.6% and 3.7%, respectively, in comparison to the most recent YOLOv8 algorithm. The mAP and recall rate on the public data set VOC2007 increased by 2.9% and 2.2%, respectively.
京公网安备11010802044758号