AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.9 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Frustum-based camera-radar fusion for 3D object detection in agricultural scenes

Lili YANGXiao GUOZi´ang LICaicong WU( )
College of Information and Electrical Engineering, China Agricultural University, Beijing 100083, China
Show Author Information

Abstract

A perception system is one of the most important components for the autonomous driving of agricultural machinery. However, there are only a few perception datasets specifically designed for agricultural scenarios, due to their difference from the typical urban scenarios in previous studies. In contrast to the urban examples, the agricultural applications can suffer from harsh working circumstances. It is often required for the perception sensors and algorithms. In this study, a low-cost perception system was proposed for the two-stage detection using millimeter-wave radar and a monocular camera. 3D object detection was then performed on the autonomous driving of the agricultural machinery under agricultural scenarios. Firstly, a multimodal perception dataset of the agricultural scenes was constructed to incorporate the LiDAR (light detection and ranging), INS (inertial navigation system), camera, and millimeter-wave radar data with a hardware-level data synchronization and target-level data annotation. Then the middle fusion strategy was used to build a neural network model, known as CFPNet. Preliminary detection of the target was also implemented with the improved network of the center point detection. Furthermore, the radar point cloud features were extracted from the frustum region of interest to supplement the image features. Finally, the preliminary detection information and radar feature were combined to perform a secondary detection. The 3D object attributes (depth, direction, and velocity) were regressed concurrently. The results show that the mAP (mean average precision) of the CFPNet on the self-built multimodal dataset of the agricultural perception was 86.5%, which was 5.5 percentage points higher than the baseline, and the mATE (mean average translation error) was 0.197 m lower than the baseline. An additional experiment on small object detection was conducted to verify the effectiveness of the CFPNet. The better performance was achieved in a recall rate of 1 for the selected small objects, which was 0.3 higher than before the improvement, indicating the better performance of the detection. Deployment experiments were conducted to test the applicability of the CFPNet. A frame rate of 7.4 frames per second was 211% of the baseline in the low-computing agricultural scenarios. Experiments on the public datasets were conducted to test the CFPNet in the rest scenarios. The favorable performance was achieved on the NuScenes public dataset, with the mATE, mASE, and mAVE of 0.792 m, 0.236, and 0.52 m/s, respectively. Since the CFPNet was specifically designed for monocular cameras, its mAP lagged behind. Furthermore, the CFPNet can directly provide the speed information of the target without the preceding and following frames. This finding can provide a feasible solution and technical support for the 3D object detection in agricultural scenarios, especially with low computing power.

CLC number: S232.3 Document code: A Article ID: 1002-6819(2026)-04-0024-09

References

【1】
【1】
 
 
Transactions of the Chinese Society of Agricultural Engineering
Pages 24-32

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
YANG L, GUO X, LI Z, et al. Frustum-based camera-radar fusion for 3D object detection in agricultural scenes. Transactions of the Chinese Society of Agricultural Engineering, 2026, 42(4): 24-32. https://doi.org/10.11975/j.issn.1002-6819.202508160

258

Views

1

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 19 August 2025
Revised: 19 November 2025
Published: 28 February 2026
© Chinese Society of Agricultural Engineering 2026