Publications
Sort:
Open Access Research Article Issue
Few-shot driven construction method of a large-scale light-trapped insect annotation data based on vision foundation models and self-supervised learning
Journal of Integrative Agriculture (JIA) 2026, 25(7): 2915-2935
Published: 26 August 2025
Abstract PDF (73.6 MB) Collect
Downloads:0

The intelligent pest-monitoring light trap based on machine vision employs specific light spectra to attract pests, infrared heating to eliminate pests, and artificial intelligence models to recognize and count them. Achieving optimal model performance requires a high-quality insect annotated dataset. However, traditional manual annotation is expert-dependent, time-consuming, and inefficient for large-scale multi-class insect labeling. This study establishes an efficient, few-shot learning approach to construct a large-scale light-trapped insect dataset through a two-stage annotation framework: detection followed by classification. Specifically, a MLTIDD addresses scale and receptive field disparities between large and tiny insects. Based on a fine-tuned Grounding DINO, SAM and SAHI are integrated to detect insects at multiple scales. Subsequently, InsectSSRL, an iBOT-based self-supervised method, learns robust insect feature representations from the extensive set of unlabeled insect sub-images detected by MLTIDD. It enhances feature extraction capability for insect sub-images through three proxy tasks. This feature extractor supports a classification model to pre-classify insect sub-images. Following expert correction, labels are traced back to original images to complete annotation work for the light-trapped insect dataset. Experimental results demonstrate that under limited samples, MLTIDD achieved 79.6% average precision (AP)50–95 and 90.8% average recall (AR), surpassing DINO by 7.0 and 4.7 percentage points. InsectSSRL attained 85.87% top-1 accuracy in k-NN evaluation. In few-shot classification, Swin-T pre-trained with InsectSSRL and fine-tuned on 5% of InsectID achieved 80.35% accuracy, exceeding iBOT by 2.08 and COCO-based transfer learning by 11.3 percentage points. The proposed pipeline improved mAP50–95 by 10.91 and AR by 8.26 percentage points compared to DINO and iBOT, while reducing expert annotation time by approximately 80% relative to manual labeling.

Total 1