AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (73.6 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Few-shot driven construction method of a large-scale light-trapped insect annotation data based on vision foundation models and self-supervised learning

Yanchen You1Zelin Feng2,3Zhe Wang3Lingyi Li3Ju Luo4Jun Lü1Haowen Zhang3Baojun Yang4Shuhua Liu4Qing Yao3( )
School of Information Science and Engineering, Zhejiang Sci-Tech University, Hangzhou 310018, China
School of Information and Control, Keyi College of Zhejiang Sci-Tech University, Hangzhou 310018, China
School of Computer Science and Technology, Zhejiang Sci-Tech University, Hangzhou 310018, China
State Key Laboratory of Rice Biology and Breeding, China National Rice Research Institute, Hangzhou 311401, China
Show Author Information

Highlights

• Multi-scale light-trapped insect-DINO detector (MLTIDD) with segment anything model (SAM) and slicing-aided hyper inference (SAHI) effectively detects tiny insects in light-trapped images.

• MLTIDD demonstrates robust generalization across diverse light-trapped scenarios.

• InsectSSRL with multiple proxy tasks learns robust insect feature representations.

• Vision transformer (ViT) trained with InsectSSRL exhibits exceptional few-shot learning capability.

• Proposed data construction method reduces expert annotation time by 80% while maintaining precision.

Abstract

The intelligent pest-monitoring light trap based on machine vision employs specific light spectra to attract pests, infrared heating to eliminate pests, and artificial intelligence models to recognize and count them. Achieving optimal model performance requires a high-quality insect annotated dataset. However, traditional manual annotation is expert-dependent, time-consuming, and inefficient for large-scale multi-class insect labeling. This study establishes an efficient, few-shot learning approach to construct a large-scale light-trapped insect dataset through a two-stage annotation framework: detection followed by classification. Specifically, a MLTIDD addresses scale and receptive field disparities between large and tiny insects. Based on a fine-tuned Grounding DINO, SAM and SAHI are integrated to detect insects at multiple scales. Subsequently, InsectSSRL, an iBOT-based self-supervised method, learns robust insect feature representations from the extensive set of unlabeled insect sub-images detected by MLTIDD. It enhances feature extraction capability for insect sub-images through three proxy tasks. This feature extractor supports a classification model to pre-classify insect sub-images. Following expert correction, labels are traced back to original images to complete annotation work for the light-trapped insect dataset. Experimental results demonstrate that under limited samples, MLTIDD achieved 79.6% average precision (AP)50–95 and 90.8% average recall (AR), surpassing DINO by 7.0 and 4.7 percentage points. InsectSSRL attained 85.87% top-1 accuracy in k-NN evaluation. In few-shot classification, Swin-T pre-trained with InsectSSRL and fine-tuned on 5% of InsectID achieved 80.35% accuracy, exceeding iBOT by 2.08 and COCO-based transfer learning by 11.3 percentage points. The proposed pipeline improved mAP50–95 by 10.91 and AR by 8.26 percentage points compared to DINO and iBOT, while reducing expert annotation time by approximately 80% relative to manual labeling.

References

【1】
【1】
 
 
Journal of Integrative Agriculture (JIA)
Pages 2915-2935

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
You Y, Feng Z, Wang Z, et al. Few-shot driven construction method of a large-scale light-trapped insect annotation data based on vision foundation models and self-supervised learning. Journal of Integrative Agriculture (JIA), 2026, 25(7): 2915-2935. https://doi.org/10.1016/j.jia.2025.08.020

185

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 14 April 2025
Revised: 24 June 2025
Accepted: 01 July 2025
Published: 26 August 2025
© 2026 CAAS.

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). Peer review under responsibility of Editorial Board of Journal of Integrative Agriculture.