Attempts to deploy computer vision in agricultural tasks often suffer from a shortage of annotated data. One strategy to alleviate the impact of limited data is Self-Supervised Learning (SSL), which involves pre-training a model on a pretext task that utilizes automatically generated annotations. The primary objective of this study is to leverage a multi-camera view dataset of cotton boll images for contrastive learning in order to enable phenotyping tasks with minimal data annotation. This dataset was collected in the field using six camera views. The efficacy of two contrastive learning frameworks (SimCLR and MoCo) in producing representations when positive examples originate from different cameras was investigated, and a comprehensive study of how the camera positions affect performance was conducted. After self-supervised pre-training, linear evaluation and semi-supervised learning experiments were performed on boll detection and plot status downstream tasks. In general, using multiple camera views with SimCLR and MoCo improves cotton boll detection mean average precision by 14% compared to vanilla SimCLR and MoCo. Through careful investigation using synthetic data, it was determined that relative camera poses with an intermediate amount of overlap seem more likely to perform well. Neither MoCo nor SimCLR was consistently superior to the other in this context. The representations embed meaningful features about the cotton plants, such as overall boll density, but also less meaningful ones, such as lighting variations. This technique could potentially accelerate the development of phenotyping algorithms based on data collected from field robots.
- Article type
- Year
- Co-author
Open Access
Research Article
Issue
Open Access
Research Article
Issue
Unmanned aerial systems (UAS) are reliable tools for field phenotyping, enabling rapid, large-scale, and costeffective data collection to support breeding programs. However, many UAS-based approaches rely on manual data processing, limiting scalability and efficiency. This study presents a fully automated pipeline for highthroughput phenotyping (HTP) of peanut crop architectural traits, including canopy height (CH), growth habit (GH), and mainstem prominence (MP) by integrating UAS imagery, a vision foundation model—Segment Anything Model (SAM), and convolutional neural networks (CNN). SAM auto-mask generator mode was used to identify field extent and orientation, while SAM interactive mode enabled individual plot segmentation using auto-generated point prompts. Terrain points automatically sampled near each plot were used to model the ground surface and compute the canopy height model, allowing CH estimations at the plot level. CH estimations showed strong agreement with manual measurements (R2 = 0.78, RMSE = 3 cm, MAPE = 10 %). For MP and GH estimation, three pre-trained CNN models (AlexNet, ResNet18, and EfficientNet-B0) were evaluated, with AlexNet achieving the highest accuracy (89 % for GH, 83 % for MP). To assess the feasibility of using these HTP-derived estimations in plant breeding, quantitative trait loci (QTL) analysis was performed, identifying major-effect loci associated with these traits. The results were consistent with conventional QTL mapping methods, demonstrating that UAS-based phenotyping provides reliable trait data for genetic studies in peanut breeding. Overall, our deep learning-based data processing workflow minimizes manual efforts, providing an efficient and scalable approach that can accelerate genetic studies and trait selection in large-scale breeding programs.
京公网安备11010802044758号