Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Aiming at the problems of low target pixels and intricate background in small target detection in infrared scenes, a target detection model based on multi-scale feature extraction with YOLOv8 was proposed. Firstly, all downsampling convolutions in the network were replaced with the Haar wavelet downsampling (HWD) module to better preserve fine-grained details in infrared imagery during downsampling. Secondly, the spatial pyramid pooling-fast (SPPF) module was improved by introducing separable convolutions, which expanded the receptive field in both horizontal and vertical directions, enabling more comprehensive spatial information capture. Furthermore, a novel C2f_CDWR module was designed using dilated convolutions with varying dilation rates to achieve adaptive feature extraction across multiple receptive fields, thus enhancing detection performance for objects of different sizes. Finally, to improve localization accuracy, the original CIoU loss in YOLOv8 was replaced with Inner-SIoU, which effectively improved bounding box regression accuracy and significantly boosted the model’s capability in detecting small infrared targets. The experimental evaluation on the HIT-UAV dataset shows that the precision of the enhanced YOLOv8 model is 90.5%, the recall rate is 75.9%, and the mean average precision is 85.7%. In terms of infrared target detection, its performance was significantly better than that of the baseline YOLOv8 model and other benchmark models.
The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits use, distribution and reproduction in any medium, provided the original work is properly cited.
Comments on this article