Accurate and high-throughput detection of rapeseed pods has been a key prerequisite for phenotypic analysis and genetic breeding. However, the axis-aligned object detection cannot fully meet the slender morphology, unpredictable spatial orientation, and high-density overlap of pods in natural field environments. Severe background interference and feature fragmentation have often limited to maintain bounding box integrity for the elongated structures, leading to the labor-intensive manual post-processing. In this study, a rotating object detection framework (SwinPodDet) was proposed to be designed specifically for the slender targets. A non-destructive pipeline was also provided to precisely localize and quantify the individual pods directly from high-resolution field imagery, thereby bypassing the constraints of axis-aligned detectors. Several key innovations were introduced for the complex geometry and topology. Firstly, a special backbone (R-SwinTransformer) was integrated with an Adaptive Shifting Window Multi-Head Self-Attention (ASW-MSA) mechanism. The window aspect ratio and shift offsets of the backbone were dynamically adjusted to effectively capture the long-range dependencies of elongated pods. The backbone was used to bridge the structural gap between local surface textures and global symmetry. Secondly, an Elongated Feature Enhancer (EFE) module was embedded to reinforce the directional sensitivity. The anisotropic depthwise separable convolutions were employed with the asymmetric 1×11 and 11×1 kernels to combine with a dual-attention mechanism. Features were selectively amplified along the pod's principal axis to suppress the orthogonal environmental noise and branch interference. Thirdly, a Multi-Scale Context Channel Attention (MSCAA) module was integrated into the feature-pyramid neck. Four parallel heterogeneous branches were utilized from local average pooling to dilated separable convolutions. MSCAA was adaptively fused with multi-scale contextual information using learnable weights. Missed detections were significantly mitigated in dense and overlapping clusters, where boundary definitions were often blurred. The model was trained and validated on the newly constructed Rotated Bounding Box Rapeseed Pod Dataset (RBRD), which was a benchmark with the 8 505 manually annotated rotating boxes. The crop was also captured over multiple growth stages. Experimental results demonstrate that the SwinPodDet was achieved in a precision (P) of 98.50%, a recall (R) of 80.76%, and a mean average precision (mAP50) of 81.74%. Furthermore, there were improvements of 2.62, 0.12, and 0.18 percentage points over the same metrics, respectively, compared with the rotating object detection network with a standard Swin-Transformer backbone. The high computational efficiency was also maintained with a parameter count of 53.32 M and an inference speed of 19.0 frames per second. The framework was achieved in an optimal Pareto balance between accuracy and deployment costs. Ablation studies and visualization analysis confirm that the "feature fracture" was effectively resolved in the high-overlap scenarios, with an impressive 95.51% counting accuracy at the whole-plant level. This performance at different maturation stages—from green-succulent to yellow-gray shriveled phases—demonstrated the superior adaptability to varying field conditions. The rapeseed pod detection was realized to treat the challenges, namely dense overlap, extreme elongation, and orientation variability. The reliable, scalable, and end-to-end tool was obtained for high-throughput phenotyping. This finding can also provide a solid algorithmic foundation for the yield estimation and large-scale agronomic monitoring in precision agriculture.
Publications
- Article type
- Year
- Co-author
Year
Issue
Transactions of the Chinese Society of Agricultural Engineering 2026, 42(7): 226-238
Published: 15 April 2026
Downloads:3
Total 1
京公网安备11010802044758号