Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Plant height is one of the key phenotypic traits to reflect the crop growth for the yield prediction in precision agriculture. It is often required to monitor the wheat plant height in the field during crop breeding and cultivation. However, the conventional measurements of plant height in the field cannot fully meet the high scalability and consistency in large-scale agricultural production, due to the labor-intensive, time-consuming, and subjective errors. In this study, an estimation framework of wheat height was proposed to integrate the semantic segmentation over the entire growth stages. UAV RGB images were acquired from multiple perspectives. Spatiotemporal patterns were combined with structure from motion (SfM) 3D reconstruction to generate digital surface models (DSM) and digital terrain models (DTM). Then the crop height model (CHM) was derived to subtract the DSM from the DTM. Meanwhile, an improved SegFormer model was employed for semantic segmentation to accurately extract wheat canopy regions from field images. Background noise was effectively eliminated, such as the soil and non-vegetation elements. According to the segmentation masks and CHM, a height inversion was established to realize the accurate and efficient estimation of wheat canopy height. Specifically, the masks were used to isolate canopy regions in the CHM. Its canopy height was represented by the 95th percentile of height values within each connected vegetation cluster. The average wheat height was derived at the plot level over each growth stage, enabling reliable phenotypic and temporal monitoring. A parallel structure was combined with a CNN and a Transformer semantic branch in the encoder of the segmentation model. The synergistic representation was then enhanced by the local texture features and global contextual information. The CNN branch was employed to capture subtle edge structures and local texture variations, particularly when the canopies were sparse and fragmented during early growth stages. In contrast, the Transformer branch was used to encode the long-range dependencies and semantic context, leading to the robust representation of large-scale canopy structures. In the decoder, progressive upsampling with skip connection was combined with a feature fusion module to effectively integrate multi-scale features. Thereby, the spatial information was preserved to refine the boundary features. An Aggregation Layer was introduced as a feature fusion module to effectively combine heterogeneous features from the dual-branch encoder and multi-scale decoder. Point-wise multiplication and addition were used to enhance feature correlations between local and global representations. Additionally, the convolutional refinement, normalization, and channel recalibration were incorporated to improve the stability and semantic consistency of the fused output. Experimental results demonstrate that the better performance was achieved in the mean intersection over union (mIoU), mean pixel accuracy (mPA), and pixel accuracy (PA) values of 80.92%, 89.42%, and 90.07%, respectively, outperforming the original SegFormer model. The canopy heights were estimated with a strong correlation after field measurements, with a coefficient of determination (R2) of 0.985, a root mean square error (RMSE) of 0.73cm, and a relative RMSE (rRMSE) of 2.41%. In addition, the temporal dynamics of wheat growth shared consistent height accumulation over the different stages. The improved model efficiently and accurately retrieved the spatial distribution and plant height information of wheat canopies. The finding can provide reliable technical support to dynamically monitor the winter wheat growth and field phenotyping in precision agriculture.
Comments on this article