Publications
Sort:
Open Access Issue
Lightweight fresh tea leaf recognition method based on improved YOLOv5s
Journal of Intelligent Agricultural Mechanization 2025, 6(1): 1-14
Published: 15 February 2025
Abstract PDF (23.2 MB) Collect
Downloads:0

The classification and recognition of tea buds represents a crucial aspect of renowned tea production. In view of the problems of large model size,large computational complexity and inability to distinguish the picking morphology of the current tea bud recognition algorithm,this study proposes an enhanced fresh tea leaf recognition model(YOLOv5s-SPCS)based on YOLOv5s as the foundational model. Firstly,images of fresh tea leaves were collected in both laboratory and natural environments to create a dataset of fresh tea leaves. This was done through offline and online collection of images in multiple scenarios,with the resulting images divided into a training set and a test set. Secondly,the Shuffle Block module was constructed based on the ShuffleNetV2 idea for replacing the convolution module in YOLOv5s backbone network,which reduced the number of model parameters and the amount of computation while increasing the speed of feature extraction. Subsequently,the Partial Convolution structure,PConv and SimAM were incorporated into the neck network to construct the C3-PCS module,replacing the original C3 structure which further reduced the model computational redundancy and memory access,while improving the recognition accuracy with a minimal increase in the number of parameters. Finally,the SIoU bounding box loss function was employed to enhance the convergence velocity and precision of the prediction frame. In addition to accelerating the convergence of the model prediction frame regression,the use of this loss function also generates more accurately positioned prediction frames. The experimental results demonstrate that the enhanced YOLOv5s-SPCS model exhibits 14%,14% and 16% of the YOLOv5s model in terms of the number of parameters,computational volume and weight file. The size of the model is,respectively,with an accuracy of 81.8% and a mean average precision(mAP)of 82.4% for the fresh tea image recognition,which is 2.7% more accurate than the original model. The accuracy was enhanced by 2.7 percentage points,while the mean average precision of mAP remained unaltered. Furthermore,the overall performance of the enhanced YOLOv5s-SPCS model is superior to that of the prevailing target detection models,including Faster R-CNN,SSD,YOLOv3,and YOLOv4. This study offers a valuable technical foundation for fresh tea leaves recognition classification and subsequent mobile deployment.

Issue
Detecting the key points of tractor drivers under complex environments using improved YOLO-Pose
Transactions of the Chinese Society of Agricultural Engineering 2023, 39(16): 139-149
Published: 30 August 2023
Abstract PDF (2.9 MB) Collect
Downloads:0

Key point leakage and misdetection have posed a great challenge on the recognition of tractor driver, due to the light, background, and occlusion in the complex operating environment of farmland. In this study, a joint driver-key point detection was proposed using improved YOLO-Pose. Firstly, Swin Transformer encoder was introduced in the top layer C3 module of the backbone network CSPDarkNet53. Among them, the encoder window size was set as 8, and the number of self-attention computation heads was 16. Swin Transformer encoder was used the self-attention of shifted windows (SW-MSA) computation to learn the cross-window interactions. The masking mechanism was utilized to isolate the invalid information exchange between pixels in non-adjacent regions in the original feature map. The better performance was achieved in the dense prediction and high-resolution vision, compared with the traditional ViT architecture. The improved model was obtained to effectively capture the global dependencies with the high computational efficiency. The global modelling capability was then improved the detection efficiency of key point under the occlusion condition. Secondly, RepGFPN, an efficient layer aggregation network with hopping structure and cross-scale connectivity, was adopted as the neck network, where the P6 detection layer was additionally added into the multi-scale output of the backbone network. CspStage module was adopted with the reparameterized ideas and layer aggregation connectivity to fuse the high-level semantic information and the low-layer spatial information, in order to enhance the model multi-scale detection. Thirdly, the pyramid convolution was introduced with 4-layer pyramid structure to replace the standard 3×3 convolution, in order to further optimize the neck network. The bottom-up layer-by-layer increasing convolution kernel was utilized to adaptively adjust the receptive field in the pyramid convolution. The number of model parameters was reduced to effectively capture the feature information of different layers. Finally, the decoupling head of key point was optimized to embed the coordinate attention mechanism, and then encode the horizontal and vertical position information into the channel attention. The network was obtained to acquire the cross-channel information, and then capture the direction-aware and position-sensitive information. The better capture performance was also achieved in the positional relationship between key points in the prediction process, indicating the high detection accuracy of the key points in the complex environments. The experimental results show that the improved model shared a mean average precision (mean average precision, mAP0.5) of 89.59%, when the Loks (object keypoint similarity) threshold was taken as 0.5, and a mAP of 0.5:0.95 (Loks thresholds were taken as 0.5, 0.55,..., 0.95, when the mean average precision) was 62.58%, which was 4.24 and 4.15 percentage points higher than the baseline model, respectively, and the average detection time of a single image was 21.9 ms. Furthermore, the mAP0.5 was improved by 7.94, 5.27, and 2.66 percentage points, and the model size was reduced by 257.5, 8.2, and 9.3 M, respectively, compared with the current mainstream networks of key point detection, such as Hourglass, HRNet-W32, and DEKR. The improved detection of key point presented the high detection accuracy and inference speed in complex scenes, especially in the case of the driver's presence of self-obscuring and other-object-obscuring. The finding can provide a strong theoretical basis for the driver behavior recognition and state monitoring in farmland operation environment.

Total 2