AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.7 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Just Accepted

VQA-G Annotator: A Self-Improving Agentic Framework for High-Fidelity Grounded Dataset Synthesis

Rongsheng Hu1Yuan Liu1Runwei Guan2Yicheng Di1Jiayu Bao1Yi Jin1Hongjian Shi3Yuan Hong4Lin Yu5Ruhui Ma3( )Rajkumar Buyya6( )

1 School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi, 214122, China

2 Thrust of Artificial Intelligence, Hong Kong University of Science and Technology (Guangzhou), Guangzhou, 511453, China

3 School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China

4 UGO-AI Intelligent Technology (Shanghai) Co.,Ltd, Shanghai, 201206, China

5 Shanghai Jinqiao Intelligent Connected Vehicle Development Co.,Ltd, China (Shanghai) Pilot Free Trade Zone, 200131, China

6 University of Melbourne, Victoria, 3010, Australia

Show Author Information

Abstract

As artificial intelligence systems increasingly rely on multi-source perception and cross-modal learning for human-like visual understanding, their outputs should be both semantically valid and spatially grounded. However, large-scale vision-language benchmarks with precise grounding remain costly to construct, while existing synthetic pipelines often suffer from distribution drift, error propagation, and weak spatial verification. We introduce VQA-G Annotator, a self-improving agentic framework that combines distribution-aware planning, multi-agent orchestration, and spatial reasoning verification in a closed-loop generation, evaluation, and refinement process. Experiments across four benchmark sources show that VQA-G Annotator achieves an average VQAScore of 0.88 for semantic alignment and an Acc@0.5 of 0.57 for spatial grounding, comparing favorably with strong automatic baselines. Downstream and human evaluations further support the utility of the synthesized data 18 for scalable and trustworthy grounded visual understanding.

References

【1】
【1】
 
 
Tsinghua Science and Technology

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Hu R, Liu Y, Guan R, et al. VQA-G Annotator: A Self-Improving Agentic Framework for High-Fidelity Grounded Dataset Synthesis. Tsinghua Science and Technology, 2026, https://doi.org/10.26599/TST.2026.9010076

77

Views

3

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 31 December 2025
Revised: 13 April 2026
Accepted: 21 July 2026
Available online: 21 July 2026

© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).