Publications
Sort:
Issue
Method for estimation of bagged grape yield using a self-correcting NMS-ByteTrack
Transactions of the Chinese Society of Agricultural Engineering 2023, 39(13): 182-190
Published: 15 July 2023
Abstract PDF (2.5 MB) Collect
Downloads:2

Overlapping occlusion has seriously limited the yield estimation of bagged grapes in recent years, due to the ever-increasing grape volume after bagging and the large surface area of grape leaves. The unstable speed of manual video shooting can also lead to the loss of bagging grape targets. In this study, a yield estimation was proposed for the bagged grape using self-correcting Non-Maximum Suppression (NMS)-ByteTrack. Firstly, the bagged grapes were detected in the video using object detection (YOLOv5s). The NMS operation was also post-positioned to the tracking stage to retain the fruit detection boxes that filtered under the occlusion. Specifically, the detection boxes of bagged grapes were detected by object detection (YOLOv5s), whereas, the prediction boxes of bagged grapes were predicted to calculate the intersection over union using the Kalman filter. If the intersection over the union was less than the given threshold, the detection box was filtered to reserve the detection boxes closer to the predicted position. Then, the final detection box was obtained using the NMS operation. Secondly, the camera motion compensation and improved Kalman filter algorithm were added to automatically correct the position of the fruit prediction boxes and track them using ByteTrack. Specifically, the camera motion compensation was used to first extract the background key points in the bagged grape pictures of the previous frame and the current frame except for the tracking target. The sparse optical flow was used to match the extracted background in the key points of the bagged grapes. Then the affine transformation matrix of background motion was calculated by the RANSAC algorithm, and the Kalman filter was utilized to predict the bagged grape prediction boxes of the current frame. Finally, the affine transformation matrix was obtained to convert the prediction boxes in the coordinate system from the previous to the current frame, in order to realize the position information correction for the prediction boxes of the current frame. In addition, the improved Kalman filter algorithm can be used to change the aspect ratio in the state vector, and then to express the position of the bagged grape tracking boxes in the Kalman filter directly as the width, in order to more accurately estimate the position of the bagged grape tracking boxes. As such, a line-counting strategy was proposed to automatically count the bagged grapes. The strategy was also to set the counting line in the middle of the video, and also retain the idle frames without the bagged grapes in the first few seconds of the video. Furthermore, the bagged grapes with the same ID were only counted once to avoid repeated collision line counting. The successful collision between the center point of the tracking box of bagged grapes and the counting line was achieved in the automatic counting of bagged grapes. The dataset was collected from the Paidengte Agricultural Science and Technology Demonstration Park, Bishan District, Chongqing of China. The mobile phone cameras of Redmi K40 and OPPO Reno6pro+ were selected to capture the pictures of the same bagged grape at 8:00, 12:00, and 18:00 from different angles, such as frontal, sideways, and overhead shots. The total shooting time was about 6 h, the shooting height was about 1.5 m from the ground, and the shooting route was line by line. A total of 500 images and six videos were obtained for the bagged grapes. Among them, 500 images were expanded to 2000 images after saturation, brightness enhancement, brightness reduction, and mirror operations. Then, 2000 images were randomly divided into the training and validation sets in the ratio of 8:2, where six valid videos were used as the test set. Relevant experiments were conducted using this dataset. The experimental results showed that the yield estimation achieved a significant improvement in the tracking performance of bagged grapes using self-correcting NMS-ByteTrack. The multi-object tracking accuracy and precision, as well as the identification F1-score, were 64.6%, 82.4% and 80.8%, respectively, which increased by 1.7, 1.0 and 4.1 percentage points, respectively, compared with the ByteTrack. The number of ID switch was reduced by 3 times. Then, the average counting accuracy reached 82.8% in terms of counting performance, compared with manual counting. In addition, a comparison was also made with the five tracking methods. A better tracking performance was achieved, compared with the rest. The applicability of this estimation was also verified in tracking and counting bagged grapes. Therefore, the yield estimation can be expected to effectively promote the tracking and counting of bagged grapes in real scenarios using self-correcting NMS-ByteTrack. More accurate yield estimation of bagged grapes was also achieved in this field.

Issue
Real-time detecting and counting dual-association bagged grape clusters using EMO-YOLOv5s
Transactions of the Chinese Society of Agricultural Engineering 2025, 41(12): 161-171
Published: 30 June 2025
Abstract PDF (3.8 MB) Collect
Downloads:1

An accurate real-time counting is vital for the bagged grape clusters, in order to ensure the subsequent estimation of the orchard yield. However, some challenges are still remained on the current real-time fruit counting. The tracking target can be lost in the bagged grape clusters, due to the dense distribution, occlusion and unstable camera movement. In this study, a real-time detection and counting were combined for the dual-association bagged grape clusters using EMO-YOLOv5s. Dual-association tracking was used as the BIoU and Euclidean distance. The rectangular region counting was also selected after target tracking. Firstly, the efficient model (EMO) was introduced as the backbone network of YOLOv5s in the detection stage. The fewer parameters were sufficiently utilized by the window multi-head self-attention (W-MHSA) mechanism in the Swin Transformer and the depthwise separable convolution (DSConv). EMO-YOLOv5s was improved the inference speed. Secondly, a dual-association was realized using ByteTrack and buffered intersection over union (BIoU) in the tracking stage. Euclidean distance was proposed to solve the target loss in the bagged grape clusters tracking. The association-matching performance of the tracking stage was enhanced to conduct twice associations between the bagged grape clusters detection and the prediction boxes. Finally, the rectangular region was designed to improve the counting accuracy of the bagged grape clusters in the counting stage. The counting increased the probability of the effectively counting bagged grape clusters. The automatic counting of the bagged grape clusters was realized to enlarge the countable range of the fruits. The experimental dataset was collected from the Agricultural Science and Technology Demonstration Park in Bishan District, Chongqing, China. The coordinate longitude, latitude, and altitude of the demonstration park were 106.221°E, 29.753°N, and 353 m, respectively. The grapes were planted line by line, with the equal row spacing, and the relatively uniform distribution of the bagged grape clusters. OPPO Reno6pro+ and Redmi K40 mobile phones were used to capture the images of the bagged grape clusters. Four angles were selected as the front, top, upward, and side in three periods of 08:00, 12:00, and 18:00 on July 22, 2022. The growth of the bagged grape clusters was then evaluated at the different angles and periods. The shooting time was about 6 h, where the height was 1.0-1.8 m from the ground, and the horizontal distance was 0.1-1.0 m. The 500 original images of the bagged grape clusters were obtained to capture line by line, with the image resolutions of 4 000×3 000 pixels and 4 096×3 072 pixels. There were 6 valid videos of the bagged grape clusters. The video resolution was 1 920×1 080 pixels, where the video format was MP4, the video frame rate was 30 frames per second and the video time was about 20 s. In addition, 200 images of the bagged grape clusters were also taken to supplement the original dataset on September 1, 2023, from 09:00 to 11:00, with an image resolution of 4 096×3 072 pixels. The experiments were conducted on the self-built dataset of the bagged grape clusters. The results showed that: 1) The parameters and floating-point operations decreased by 38.6% and 39.0% in the performance of the detection, respectively, compared with the YOLOv5s. Meanwhile, the average precision and detection speed were achieved by 96.5% and 77 frames per second, respectively. 2) In the tracking performance, the higher order tracking accuracy, multiple-object tracking accuracy, and identification F1-score were 58.6%, 64.7% and 80.0%, respectively, which were 3.6, 4.1, and 6.0 percentage point higher than ByteTrack. 3) In the counting performance, the average counting precision was achieved by 93.1%. Meanwhile, the mean absolute error was 1.3%. As such, the improved model was effectively solved the problems of the tracking and counting bagged grape clusters. The finding can provide a reliable basis to predict the orchard yield.

Total 2