Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Non-Orthogonal Multiple Access (NOMA) assisted Unmanned Aerial Vehicle (UAV) communication is becoming a promising technique for future B5G/6G networks. However, the security of the NOMA-UAV networks remains critical challenges due to the shared wireless spectrum and Line-of-Sight (LoS) channel. This paper formulates a joint UAV trajectory design and power allocation problem with the aid of the ground jammer to maximize the sum secrecy rate. First, the joint optimization problem is modeled as a Markov Decision Process (MDP). Then, the Deep Reinforcement Learning (DRL) method is utilized to search the optimal policy from the continuous action space. In order to accelerate the sample accumulation, the Asynchronous Advantage Actor-Critic (A3C) scheme with multiple workers is proposed, which reformulates the action and reward to acquire complete update duration. Simulation results demonstrate that the A3C-based scheme outperforms the baseline schemes in term of the secrecy rate and stability.
This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Comments on this article