Publications
Sort:
Issue
Method for estimation of bagged grape yield using a self-correcting NMS-ByteTrack
Transactions of the Chinese Society of Agricultural Engineering 2023, 39(13): 182-190
Published: 15 July 2023
Abstract PDF (2.5 MB) Collect
Downloads:2

Overlapping occlusion has seriously limited the yield estimation of bagged grapes in recent years, due to the ever-increasing grape volume after bagging and the large surface area of grape leaves. The unstable speed of manual video shooting can also lead to the loss of bagging grape targets. In this study, a yield estimation was proposed for the bagged grape using self-correcting Non-Maximum Suppression (NMS)-ByteTrack. Firstly, the bagged grapes were detected in the video using object detection (YOLOv5s). The NMS operation was also post-positioned to the tracking stage to retain the fruit detection boxes that filtered under the occlusion. Specifically, the detection boxes of bagged grapes were detected by object detection (YOLOv5s), whereas, the prediction boxes of bagged grapes were predicted to calculate the intersection over union using the Kalman filter. If the intersection over the union was less than the given threshold, the detection box was filtered to reserve the detection boxes closer to the predicted position. Then, the final detection box was obtained using the NMS operation. Secondly, the camera motion compensation and improved Kalman filter algorithm were added to automatically correct the position of the fruit prediction boxes and track them using ByteTrack. Specifically, the camera motion compensation was used to first extract the background key points in the bagged grape pictures of the previous frame and the current frame except for the tracking target. The sparse optical flow was used to match the extracted background in the key points of the bagged grapes. Then the affine transformation matrix of background motion was calculated by the RANSAC algorithm, and the Kalman filter was utilized to predict the bagged grape prediction boxes of the current frame. Finally, the affine transformation matrix was obtained to convert the prediction boxes in the coordinate system from the previous to the current frame, in order to realize the position information correction for the prediction boxes of the current frame. In addition, the improved Kalman filter algorithm can be used to change the aspect ratio in the state vector, and then to express the position of the bagged grape tracking boxes in the Kalman filter directly as the width, in order to more accurately estimate the position of the bagged grape tracking boxes. As such, a line-counting strategy was proposed to automatically count the bagged grapes. The strategy was also to set the counting line in the middle of the video, and also retain the idle frames without the bagged grapes in the first few seconds of the video. Furthermore, the bagged grapes with the same ID were only counted once to avoid repeated collision line counting. The successful collision between the center point of the tracking box of bagged grapes and the counting line was achieved in the automatic counting of bagged grapes. The dataset was collected from the Paidengte Agricultural Science and Technology Demonstration Park, Bishan District, Chongqing of China. The mobile phone cameras of Redmi K40 and OPPO Reno6pro+ were selected to capture the pictures of the same bagged grape at 8:00, 12:00, and 18:00 from different angles, such as frontal, sideways, and overhead shots. The total shooting time was about 6 h, the shooting height was about 1.5 m from the ground, and the shooting route was line by line. A total of 500 images and six videos were obtained for the bagged grapes. Among them, 500 images were expanded to 2000 images after saturation, brightness enhancement, brightness reduction, and mirror operations. Then, 2000 images were randomly divided into the training and validation sets in the ratio of 8:2, where six valid videos were used as the test set. Relevant experiments were conducted using this dataset. The experimental results showed that the yield estimation achieved a significant improvement in the tracking performance of bagged grapes using self-correcting NMS-ByteTrack. The multi-object tracking accuracy and precision, as well as the identification F1-score, were 64.6%, 82.4% and 80.8%, respectively, which increased by 1.7, 1.0 and 4.1 percentage points, respectively, compared with the ByteTrack. The number of ID switch was reduced by 3 times. Then, the average counting accuracy reached 82.8% in terms of counting performance, compared with manual counting. In addition, a comparison was also made with the five tracking methods. A better tracking performance was achieved, compared with the rest. The applicability of this estimation was also verified in tracking and counting bagged grapes. Therefore, the yield estimation can be expected to effectively promote the tracking and counting of bagged grapes in real scenarios using self-correcting NMS-ByteTrack. More accurate yield estimation of bagged grapes was also achieved in this field.

Issue
Segmenting grape leaf diseases using multi-scale cross-fusion and boundary-aware network
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(17): 203-212
Published: 15 September 2025
Abstract PDF (2.5 MB) Collect
Downloads:2

Early identification is often required for the effective prevention and control of the grape leaf diseases. However, an accurate segmentation is limited to the varying sizes and diverse shapes of the grape leaves and their diseased areas, as well as the complex backgrounds and edge blurriness that caused by lighting interference. Moreover, the existing models can be improved the performance at the cost of the increasing model size and computational complexity. It is also demand for their effective deployment on the resource-constrained mobile devices. In this study, a multi-scale cross-fusion and boundary-aware segmentation network (MCBNet) was proposed to detect the grape leaf diseases, in order to reduce the computational costs for the high segmentation accuracy. A multi-scale cross-fusion decoder was also developed to effectively integrate the feature maps from the different scales. Multi-scale strip convolutional kernels and a cross-axis attention mechanism were utilized to capture the multi-scale global features. Additionally, a boundary-aware guidance module was introduced for the model sensitive to the boundary features. As such, the segmentation performance was enhanced on the edge-blurred diseases of the varying sizes. The experimental results show that: 1) The MCBNet exhibited the outstanding performance on the dataset of the grape leaf diseases. Specifically, in the leaf segmentation task, the MCBNet was improved Dice and IoU metrics by 0.6 and 1.1 percentage points, respectively, compared with the second-best network. In the disease segmentation task, the Dice and IoU metrics were enhanced by 1.3 and 1.9 percentage points, respectively. The HD metric was utilized to measure the accuracy of the segmentation boundary. The MCBNet outperformed the second-best network by 4.0 and 0.4 percentage points in the leaf and disease segmentation tasks, respectively. Additionally, the MCBNet was improved Dice and HD metrics by 1.3 and 1.6 percentage points, respectively, compared with the lightweight MetaSeg network. The better performance was achieved in the parameter counting of only 3.75M and a computational cost of 1.61 GFLOPs. There was the excellent balance between high segmentation accuracy and low computational cost. 2) The public PlantVillage dataset was further validated the generalization of the MCBNet. In the disease segmentation task, the MCBNet was improved Dice, IoU, Se, and Pre metrics to 85.2%, 74.2%, 83.8%, and 86.5%, respectively, compared with the second-best network. Furthermore, the MCBNet outperformed the second-best network by 4.38 percentage points in the HD metric, indicating the better performance on the blurred boundaries. 3) Visualization results also confirmed that the MCBNet was utilized to capture the disease regions of the various sizes in both self-built dataset and public datasets, significantly reducing the missed detections. Moreover, the boundary-aware guidance module of the MCBNet was greatly enhanced to process the edge details, fully validating its exceptional segmentation performance. In conclusion, the MCBNet can be expected to offer an efficient and precise solution for the grape leaf disease segmentation under the complex environments. Its lightweight design can be deployed on the resource-constrained devices. Some limitations were still remained to balance the operational efficiency and deployment requirements. A lightweight backbone network was used for the feature extraction. However, the lightweight backbone network can limit the feature extraction for the segmentation accuracy of the disease regions. Future research can be utilized to optimize the network structure for the more precise capture of the disease regions. Additionally, the model training can still rely on a large amount of the high-quality pixel-level labeled data, which is time-consuming and costly. Therefore, the weakly supervised or semi-supervised learning can be introduced to reduce the reliance on the fine-grained annotations and lower data preparation costs. Finally, the domain adaptation can be added to enhance the stability and generalization under the variable and complex environments, such as the strong lighting or partial leaf occlusion.

Issue
Real-time detecting and counting dual-association bagged grape clusters using EMO-YOLOv5s
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(12): 161-171
Published: 30 June 2025
Abstract PDF (3.8 MB) Collect
Downloads:1

An accurate real-time counting is vital for the bagged grape clusters, in order to ensure the subsequent estimation of the orchard yield. However, some challenges are still remained on the current real-time fruit counting. The tracking target can be lost in the bagged grape clusters, due to the dense distribution, occlusion and unstable camera movement. In this study, a real-time detection and counting were combined for the dual-association bagged grape clusters using EMO-YOLOv5s. Dual-association tracking was used as the BIoU and Euclidean distance. The rectangular region counting was also selected after target tracking. Firstly, the efficient model (EMO) was introduced as the backbone network of YOLOv5s in the detection stage. The fewer parameters were sufficiently utilized by the window multi-head self-attention (W-MHSA) mechanism in the Swin Transformer and the depthwise separable convolution (DSConv). EMO-YOLOv5s was improved the inference speed. Secondly, a dual-association was realized using ByteTrack and buffered intersection over union (BIoU) in the tracking stage. Euclidean distance was proposed to solve the target loss in the bagged grape clusters tracking. The association-matching performance of the tracking stage was enhanced to conduct twice associations between the bagged grape clusters detection and the prediction boxes. Finally, the rectangular region was designed to improve the counting accuracy of the bagged grape clusters in the counting stage. The counting increased the probability of the effectively counting bagged grape clusters. The automatic counting of the bagged grape clusters was realized to enlarge the countable range of the fruits. The experimental dataset was collected from the Agricultural Science and Technology Demonstration Park in Bishan District, Chongqing, China. The coordinate longitude, latitude, and altitude of the demonstration park were 106.221°E, 29.753°N, and 353 m, respectively. The grapes were planted line by line, with the equal row spacing, and the relatively uniform distribution of the bagged grape clusters. OPPO Reno6pro+ and Redmi K40 mobile phones were used to capture the images of the bagged grape clusters. Four angles were selected as the front, top, upward, and side in three periods of 08:00, 12:00, and 18:00 on July 22, 2022. The growth of the bagged grape clusters was then evaluated at the different angles and periods. The shooting time was about 6 h, where the height was 1.0-1.8 m from the ground, and the horizontal distance was 0.1-1.0 m. The 500 original images of the bagged grape clusters were obtained to capture line by line, with the image resolutions of 4 000×3 000 pixels and 4 096×3 072 pixels. There were 6 valid videos of the bagged grape clusters. The video resolution was 1 920×1 080 pixels, where the video format was MP4, the video frame rate was 30 frames per second and the video time was about 20 s. In addition, 200 images of the bagged grape clusters were also taken to supplement the original dataset on September 1, 2023, from 09:00 to 11:00, with an image resolution of 4 096×3 072 pixels. The experiments were conducted on the self-built dataset of the bagged grape clusters. The results showed that: 1) The parameters and floating-point operations decreased by 38.6% and 39.0% in the performance of the detection, respectively, compared with the YOLOv5s. Meanwhile, the average precision and detection speed were achieved by 96.5% and 77 frames per second, respectively. 2) In the tracking performance, the higher order tracking accuracy, multiple-object tracking accuracy, and identification F1-score were 58.6%, 64.7% and 80.0%, respectively, which were 3.6, 4.1, and 6.0 percentage point higher than ByteTrack. 3) In the counting performance, the average counting precision was achieved by 93.1%. Meanwhile, the mean absolute error was 1.3%. As such, the improved model was effectively solved the problems of the tracking and counting bagged grape clusters. The finding can provide a reliable basis to predict the orchard yield.

Issue
Counting bagging grape using improved YOLOv9s and adaptive Kalman filter
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(10): 195-203
Published: 30 May 2025
Abstract PDF (4.1 MB) Collect
Downloads:0

Grapes can be one type of the fruit with the widest cultivation area, the highest yield, and extremely high economic value in China. Among them, bagging techniques can be often employed to reduce the impact of pests and diseases on the grape quality during the harvest period. An accurate yield estimation can greatly contribute to the plan picking, sales, and storage, in order to reduce the economic losses caused by supply-demand mismatches. The accurate counting of bagged grapes can be required before yield estimation. However, the existing fruit counting can usually suffer from insufficient real-time detection and tracking failure, due to the occlusion of bagged grapes and unprocessed detection noise. In this study, video counting was proposed for the bagged grapes using an improved YOLOv9s and adaptive Kalman filter. Three modules were included: the improved YOLOv9s detection model, an adaptive Kalman filter tracking algorithm, and a line-drawing counting. In detection, the original RepNCSPELAN4 module in YOLOv9s was replaced with an efficient feature enhancement module (EFEM), in order to reduce the number of model parameters for the inference speed. The performance of the improved YOLOv9s model was enhanced for sufficient real-time detection. The EFEM was designed to selectively learn from the partial feature maps of the bagged grapes, thereby enabling efficient feature extraction and faster inference. The FasterNet module was specifically utilized to efficiently extract the spatial features, in order to minimize the redundant computation and memory access. A spatially enhanced attention module (SEAM) was introduced to further improve the detection performance under occlusion conditions. The SEAM was used to learn the relationship between occluded and unoccluded areas. The occluded features were predicted and compensated to thereby improve the detection accuracy of bagged grapes under full and partial occlusion. In tracking, an adaptive Kalman filter algorithm was proposed to reduce the detection noise caused by camera shake and rapid movement. The accuracy of Kalman filter trajectory prediction was promoted after tracking. Noise estimation was automatically adjusted, according to the detection confidence. A line-drawing counting was used for the real-time counting of bagged grapes; Once the center of the bagged grape was collided with a virtual counting line, the number of bagged grapes increased by one. The experimental dataset was collected from the PaiDengTe Technology Demonstration Park in Bishan District, Chongqing, China. There were 700 original images of bagged grapes and six video clips. The dataset was randomly divided into a training set of 490 images, a validation set of 140 images, and a test set of 70 images at a ratio of 7∶2∶1. The six video clips were used to test the counting performance. Some image enhancement techniques were applied to the training set during training, such as saturation adjustment, brightness variation, image mirroring, and Gaussian noise addition, thereby expanding the training set to 2100 images. The robustness and generalization of the detection model were enhanced after enhancement. Experimental results show that the best performance of the improved YOLOv9s model (ES-YOLOv9s) outperformed five other models. The highest mean average precision and recall were 96.9% and 93.1%, respectively, while there was an inference speed of 70 frames per second. Compared with the original YOLOv9s, the ES-YOLOv9s reduced the number of parameters by 29.6%, and the number of floating-point operations decreased by 10.9G, whereas, the frame rate was improved by 20 frames per second. In terms of tracking performance, the adaptive Kalman filter tracking algorithm achieved 58.6%, 63.6%, and 78.8% in the higher-order tracking accuracy, multi-object tracking accuracy, and ID harmonic mean metrics, respectively, thus representing improvements of 4.3, 2.2, and 2.5 percentage points over ByteTrack. In terms of counting performance, line-drawing counting was achieved with an average accuracy of 80.0%, compared with manual counting. In conclusion, the video counting of bagged grapes with the improved YOLOv9s and Kalman filter also demonstrated better application potential in real-time tracking and counting. The finding can provide technical support for the pre-harvest yield estimation of bagged grapes.

Total 4