AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Home iFuture Article
PDF (14.6 MB)
Collect
AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Original Article | Open Access | Just Accepted

Hi-Agent: Hierarchical vision-language agents for mobile device control

Zhe Wu1Hongjin Lu1Junliang Xing1( )Changhao Zhang1Yuxuan Li1Yin Zhu1Yuhao Yang2Yuheng Jing3Kai Li3Kun Shao2Jianye Hao2Jun Wang4Yuanchun Shi1

1 Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China

2 Huawei Noah’s Ark Lab, Huawei Technologies Co., Ltd., Beijing 100095, China

3 Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China

4 Department of Computer Science, University College London, London WC1E 6BT, UK

Show Author Information

Abstract

Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on direct state-to-action mappings, which lack structured reason-ing and planning, and thus generalize poorly to novel tasks or unseen UI layouts. We introduce Hi-Agent, a trainable hierarchical vision-language agent for mobile control, featuring a high-level reasoning model and a low-level action model that are jointly optimized. For efficient training, we reformulate multi-step decision-making as a sequence of single-step subgoals and propose a foresight advantage function, which leverages execution feedback from the low-level model to guide high-level optimization. This design alleviates the path explosion issue encountered by Group Relative Policy Optimization (GRPO) in long-horizon tasks and enables stable, critic-free joint training. Hi-Agent achieves a new State-Of-The-Art (SOTA) 87.9% task success rate on the Android-in-the-Wild (AitW) bench-mark, significantly outperforming prior methods across three paradigms: prompt-based (AppAgent: 17.7%), supervised (Filtered BC: 54.5%), and reinforcement learning-based (DigiRL: 71.9%). It also demonstrates competitive zero-shot generalization on the ScreenSpot-v2 benchmark. On the more challenging Android-World benchmark, Hi-Agent also scales effectively with larger backbones, show-ing strong adaptability in high-complexity mobile control scenarios.

References

【1】
【1】
 
 
iFuture

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Wu Z, Lu H, Xing J, et al. Hi-Agent: Hierarchical vision-language agents for mobile device control. iFuture, 2026, https://doi.org/10.26599/IF.2026.9710002

297

Views

11

Downloads

0

Crossref

Received: 15 April 2026
Revised: 02 July 2026
Accepted: 16 July 2026
Available online: 16 July 2026

© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).