Orchid flowers are frequently characterized by random bending, tilting, and mutual occlusion in the complex natural environment. These morphological irregularities and postures have caused the serious occlusion and overlapping of flower organs (sepals and petals), making it difficult to directly and accurately extract phenotypic parameters, such as length, width, and area. Furthermore, manual measurement cannot fully meet the large-scale production in recent years, due to the labor-intensive, subjective, and damage to the fragile specimens. It is often required for the precise phenotypic extraction under these unconstrained conditions using conventional computer vision. In this study, a systematic extraction was proposed using deep learning framework with the MAS-YOLO instance segmentation model and the Pix2PixHD-CA generative adversarial network. Two stages included the accurate segmentation of occluded organs and morphological restoration of incomplete organs. 1) In the instance segmentation stage, a MAS-YOLO model was constructed using the YOLO11s-seg architecture. MobileNetv4 backbone network replaced with the original ones to realize the lightweight deployment on edge devices with limited computing resources. Universal Inverted Bottleneck (UIB) blocks were utilized to significantly reduce computational redundancy for the high feature extraction. 2) Adaptive Spatial Fusion (ASF) framework was integrated to weight and fuse features from different scales for the minimum the loss of small target information. Simultaneously, a Spatial Dynamic Integration (SDI) module was introduced to improve the feature response distinction between the orchid organs and the complex background. A dataset with 520 natural images (3865 annotated instances) was used to train and validate the segmentation model. 3) In the parameter extraction stage, a Pix2PixHD-CA generation model was developed to determine the morphological deviation between the segmented occluded organs and their real flattened states. A Coordinate Attention (CA) mechanism was embedded into the generator trunk of the Pix2PixHD network. Unlike standard channel attention, the CA mechanism decomposed channel attention into two parallel 1D feature encodings, allowing the network to form joint perception in both channel and spatial coordinate dimensions. Long-range dependencies were captured to preserve precise positional information for shape reconstruction. Consequently, the mapping relationship between the "deviated organ" and the "complete flattened organ" was established using 1 500 pairs images after alignment. The images were generated to maintain high fidelity in the texture and edge trends. The results demonstrated that the superior performance was achieved in both segmentation and parameter extraction. In segmentation, the F1 score of the MAS-YOLO model increased from 0.840 (baseline) to 0.962, indicating the high accuracy to identify the occluded and bent organs. Simultaneously, the quantity of parameter was reduced from 10.08 to 8.95 M, indicating an optimal balance between segmentation accuracy and computational efficiency. The comparison after generative restoration was performed on the phenotypic parameters between the Pix2PixHD-CA extraction from flattened images and the measurements. The Coefficient of Determination (R2) reached 0.928, 0.895, and 0.937, respectively, for organ length, width, and area. The Root Mean Square Errors (RMSE) were 1.45 mm, 0.25 mm, and 14.87 mm2, respectively. The R2 values of Pix2PixHD model for length, width, and area increased by 8.79%, 1.59%, and 3.65%, respectively, while the RMSE values decreased by 3.97%, 16.66%, and 23.34%, respectively, compared with the original ones without the attention mechanism. The optimal mapping was achieved to reduce the systematic errors caused by shape distortion using CA mechanism. Furthermore, the field test was conducted on unrelated samples. The R2 values remained above 0.877 for all three phenotypic parameters, indicating the robust generalization of the model in real-world scenarios. The pipeline was also developed to reduce the interference of morphological occlusion in natural habitats. The lightweight high-precision segmentation of MAS-YOLO was effectively combined with the morphological restoration of Pix2PixHD-CA. The organ phenotypic parameters of orchid flower were accurately extracted to significantly reduce the labor intensity and subjective errors with manual measurement. The finding can provide strong technical support and high-quality data for orchid genetic breeding, ecological statistics, and evolutionary biology.
- Article type
- Year
- Co-author
Lettuce is one of the most favorite leafy vegetable in precision cultivation. It is often required to monitor the lettuce growth for real-time robotic harvesting. Nevertheless, leafy vegetables are characterized by leaves, diverse morphologies, and thin, flexible, and deformable textures. There is a more complex structure, compared with the morphologically regular objects, such as spherical fruits, umbrella-shaped mushrooms, and conical carrots. As such, plant growth and leaf expansion can lead to mutual occlusion among lettuce plants. Furthermore, the high planting density has commonly adopted in plants, leading to the inter-leaf occlusion. Additionally, height differences between individual plants can also cause upper leaves to shade lower ones. These occlusions then result in missing data in the point cloud images, seriously affecting the accurate acquisition of key phenotypic parameters. Conventional point cloud processing has mostly developed for regularly shaped crops, making it difficult to effectively reconstruct complex structures. Therefore, it is challenging to accurately detect the complete three-dimensional phenotypic parameters of lettuce from severely incomplete point cloud data. In this study, the AdaPoinTr-ER model and progressive completion were proposed for the overall completion pipeline of incomplete lettuce images in point cloud. Since the occlusion among lettuce plants occurred at edge positions, AdaPoinTr-ER model was integrated edge attention into the AdaPoinTr framework to enhance feature extraction of geometric contour. In view of the lettuce leaves with the more complex multilayer structures, AdaPoinTr-ER model was integrated residual module into AdaPoinTr to reduce feature degradation during point cloud generation for the prediction accuracy of generated point clouds. Three sequential stages were proposed to realize the progressive completion for incomplete lettuce: (1) Mask3D was employed to segment top-view lettuce clusters with mutual occlusion, thus capturing the top-view images of incomplete lettuce plants. (2) The top-down completion model trained by AdaPoinTr-ER model was used to complete the image, and the resulting images were then input into the three-dimensional completion model for training. (3) The phenotypic parameters of the completed intact lettuce were acquired after three-dimensional completion. The experimental results demonstrated that AdaPoinTr-ER model achieved the best performance in Chamfer Distance, Earth Mover's Distance, and F1-score, compared with FoldingNet, GRNet, PCN, PoinTr, and AdaPoinTr. The ablation experiments demonstrated that the AdaPoinTr-ER model achieved a Chamfer Distance of 2.88×10-4 cm, an Earth Mover's Distance of 0.19 cm, and an F1-score of 72.68%. Compared with the original AdaPoinTr model, the Chamfer Distance and Earth Mover's Distance decreased by 43.3% and 42.4%, respectively, while the F1-score improved by 9.52 percentage points. In the lettuce phenotypic analysis, the progressive completion yielded coefficients of determination (R2) of 0.933, 0.917, and 0.903 for the projected area, crown width, and plant height, respectively, with the root mean square errors (RMSE) of 8.569 cm2, 0.434 cm, and 0.591 cm, respectively. Compared with the phenotypic analysis from incomplete lettuce, the R² values increased by 45.6%, 61.4%, and 30.7%, respectively. Furthermore, the R2 values improved by 24.9%, 29.5%, and 11.1%, respectively, whereas, the RMSE was reduced by 64.8%, 56.6%, and 37.4%, respectively, compared with direct completion using only 3D reconstruction without top-view completion. Consequently, AdaPoinTr-ER model exhibited superior performance to restore both the local geometric details and the overall shape structure of lettuce point clouds. Completing occluded lettuce point clouds is a challenging task. The occluded lettuce point clouds were accurately and effectively reconstructed to significantly improve the accuracy of phenotypic parameter extraction for leafy vegetables in densely planted environments. Thereby, the finding can provide the robust support for the intelligent and precise completion of leafy vegetable species. The completion pipeline can also be integrated with real-time robotic harvesting for fully automatic operations in plants.
Acquiring the individual parameters of collective lettuce under dense scenarios can greatly contribute to environmental regulation, yield prediction, and harvest timing determination in the growth monitoring center of the plant factory. Traditional monitoring can often involve the manual measurement of geometric parameters and root removal for fresh weight determination, leading to the less comprehensive and efficient. Fortunately, non-destructive monitoring can be expected to extract the crop phenotypic parameters using machine vision and machine learning at present. However, most existing machine learning exhibited certain limitations in the application of lettuce point clouds. For instance, the majority of application scenarios have been focused on the organ segmentation of individual crops. It is still lacking in the individual segmentation of collective plants. Additionally, the extraction of crop phenotypes from point clouds can often rely mainly on manually predefined feature quantities. There is a high demand to fully explore the effective phenotypic information within the lettuce point clouds. In this study, instance segmentation was proposed to process the point cloud data of collective crops. Subsequently, deep learning of point clouds was employed to predict the fresh weight of individual crops. The collective lettuce was also taken as the research object. A depth camera was also utilized to collect the single-plane point clouds of the collective lettuce. After point cloud preprocessing, the data was then input into the instance segmentation model (Mask3D) for training. A feature backbone network was employed to extract the features from the point clouds of the collective lettuce. A Transformer decoder was utilized to process the instance queries. Point cloud features were integrated with the instance queries through the mask module. A mask was then generated for each instance. The background and lettuce point clouds were segmented to distinguish the individual lettuce. Finally, the fresh weight prediction (FWP) network was employed to predict the fresh weight of individually segmented lettuce. The feature extraction network (PointNet) was utilized to extract the features from the segmented point clouds of individual lettuce. A multilayer perceptron was also employed to regressively predict the fresh weight of the lettuce. Experimental results indicated that the segmentation and extraction of individual lettuce point clouds were successfully achieved without over- or under-detection on the point cloud dataset. When the Intersection over Union (IoU) threshold was set to 0.75, the average precision (Ap) of instance segmentation was 0.924 for collective lettuce point clouds, superior to the instance segmentation models, such as Jsnet. Furthermore, the direct prediction of lettuce fresh weight reduced the errors associated with the manual feature extraction during processing using deep learning point cloud. The coefficient of determination (R2) and the root mean squared error (RMSE) were 0.90 and 12.42 g, respectively, indicating the superior accuracy of the fresh weight prediction network. Traditional and manual feature extraction from point cloud parameters with machine learning was achieved in the maximum R² of 0.83 and the minimum RMSE of 15.06 g. Therefore, the deep learning instance segmentation and point cloud regression can be expected to estimate the fresh weight of collective lettuce, indicating exceptional better performance. The finding can also provide significant importance for the growth monitoring, yield estimation, and harvest timing determination of facility vegetables.
京公网安备11010802044758号