Non-Orthogonal Multiple Access (NOMA) assisted Unmanned Aerial Vehicle (UAV) communication is becoming a promising technique for future B5G/6G networks. However, the security of the NOMA-UAV networks remains critical challenges due to the shared wireless spectrum and Line-of-Sight (LoS) channel. This paper formulates a joint UAV trajectory design and power allocation problem with the aid of the ground jammer to maximize the sum secrecy rate. First, the joint optimization problem is modeled as a Markov Decision Process (MDP). Then, the Deep Reinforcement Learning (DRL) method is utilized to search the optimal policy from the continuous action space. In order to accelerate the sample accumulation, the Asynchronous Advantage Actor-Critic (A3C) scheme with multiple workers is proposed, which reformulates the action and reward to acquire complete update duration. Simulation results demonstrate that the A3C-based scheme outperforms the baseline schemes in term of the secrecy rate and stability.
Publications
- Article type
- Year
Year
Open Access
Issue
Chinese Journal of Aeronautics 2025, 38(10)
Published: 06 June 2025
Total 1
京公网安备11010802044758号