AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Full Length Article | Open Access

Demonstration-enhanced policy search for space multi-arm robot collaborative skill learning

Tian GAOaChengfei YUEa( )Xiaozhe JUaTao LINb
Institute of Space Science and Applied Technology, Harbin Institute of Technology, Shenzhen 518055, China
Research Center of Satellite Technology, Harbin Institute of Technology, Harbin 150001, China

Peer review under responsibility of Editorial Committee of CJA

Show Author Information

Abstract

The increasing complexity of on-orbit tasks imposes great demands on the flexible operation of space robotic arms, prompting the development of space robots from single-arm manipulation to multi-arm collaboration. In this paper, a combined approach of Learning from Demonstration (LfD) and Reinforcement Learning (RL) is proposed for space multi-arm collaborative skill learning. The combination effectively resolves the trade-off between learning efficiency and feasible solution in LfD, as well as the time-consuming pursuit of the optimal solution in RL. With the prior knowledge of LfD, space robotic arms can achieve efficient guided learning in high-dimensional state-action space. Specifically, an LfD approach with Probabilistic Movement Primitives (ProMP) is firstly utilized to encode and reproduce the demonstration actions, generating a distribution as the initialization of policy. Then in the RL stage, a Relative Entropy Policy Search (REPS) algorithm modified in continuous state-action space is employed for further policy improvement. More importantly, the learned behaviors can maintain and reflect the characteristics of demonstrations. In addition, a series of supplementary policy search mechanisms are designed to accelerate the exploration process. The effectiveness of the proposed method has been verified both theoretically and experimentally. Moreover, comparisons with state-of-the-art methods have confirmed the outperformance of the approach.

References

【1】
【1】
 
 
Chinese Journal of Aeronautics

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
GAO T, YUE C, JU X, et al. Demonstration-enhanced policy search for space multi-arm robot collaborative skill learning. Chinese Journal of Aeronautics, 2025, 38(3). https://doi.org/10.1016/j.cja.2024.08.018

722

Views

13

Crossref

5

Web of Science

10

Scopus

1

CSCD

Received: 22 February 2024
Revised: 07 April 2024
Accepted: 27 May 2024
Published: 19 August 2024
© 2024 Chinese Society of Aeronautics and Astronautics.

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).