AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

UAV Encounter Scenario Generation Method Based on Reward Shaping

Bingjie YANG1,2Xuejun ZHANG1,2( )Zhuoya LIU1,3Yue XIAO1,3
School of Electronic Information Engineering, Beihang University, Beijing 100191, China
State Key Laboratory of CNS/ATM, Beihang University, Beijing 100191, China
Beijing Key Laboratory for Network-Based Cooperative Air Traffic Management, Beihang University, Beijing 100191, China
Show Author Information

Abstract

To address the encounter scenario generation problem for validating UAV Detect-and-Avoid (DAA) systems, this paper proposes an improved PPO algorithm—RSC-PPO—which is based on reward shaping and action constraints. To guarantee the physical fidelity of the scenarios, a Markov Decision Process (MDP) based on variable-speed Dubins dynamics is established. Furthermore, an action smoothing constraint regularization term is introduced to effectively suppress drastic fluctuations in policy output, ensuring that the intruder trajectories possess physical realism. Moreover, to overcome the challenges of exploration difficulty and poor convergence in sparse reward environments for reinforcement learning, a multi-level reward shaping mechanism comprising heading accuracy guidance, circling penalty, and time pressure penalty is developed. This mechanism converts the sparse terminal encounter signal into gradual and continuous process guidance signals, thereby markedly enhancing the exploration efficiency of high-risk trajectories. Experimental results demonstrate that the trained RSC-PPO policy can efficiently and stably generate highly adversarial scenario sets, achieving a success rate exceeding 95% and realizing 100% coverage of the initial state space. The generated scenario set covers multiple typical conflict configurations, including head-on, converging, and overtaking scenarios, fully ensuring the diversity of encounters. Risk assessment based on the Effective Maneuvering Time Window (EMTW) confirms that the generated scenarios exhibit significant characteristics of high urgency. This study provides an efficient and robust scenario dataset and generation framework for the validation of safety-critical aviation systems such as DAA.

CLC number: V249.7 Article ID: 1000-565X(2026)06-0183-10

References

【1】
【1】
 
 
Journal of South China University of Technology (Natural Science Edition)
Pages 183-192

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
YANG B, ZHANG X, LIU Z, et al. UAV Encounter Scenario Generation Method Based on Reward Shaping. Journal of South China University of Technology (Natural Science Edition), 2026, 54(6): 183-192. https://doi.org/10.12141/j.issn.1000-565X.250411

9

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 28 October 2025
Published: 01 June 2026
© Journal of South China University of Technology(Natural Science Edition)