Publications
Sort:
Issue
An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios
Unmanned Systems 2025, 13(4): 987-1004
Published: 30 September 2024
Abstract Collect

Aiming at the problem of multi-agent cooperative confrontation in seize-control scenarios, we design an efficient multi-agent policy self-play (EMAP-SP) learning method. First, a multi-agent centralized policy model is constructed to command the agents to perform tasks cooperatively. Considering that the policy being trained and its historical policies usually have poor exploration capability under incomplete information in self-play trainings, the intrinsic reward mechanism based on random network distillation (RND) is introduced in the self-play learning method. In addition, we propose a multi-step on-policy deep reinforcement learning (DRL) algorithm assisted by off-policy policy evaluation (MSOAO) to learn the best response policy in the self-play. Compared with DRL algorithms commonly used in complex decision problems, MSOAO has more efficient policy evaluation capability, and efficient policy evaluation further improves the policy learning capability. The effectiveness of EMAP-SP is fully verified in MiaoSuan wargame simulation system, and the evaluation results show that EMAP-SP can learn the cooperative policy of effectively defeating the Blue side’s knowledge-based policy under incomplete information. Moreover, the evaluations results in DRL benchmark environments also show that the best response policy learning algorithm MSOAO can promote the agent to learn approximately optimal policies.

Open Access Full Length Article Issue
When LoRa meets distributed machine learning to optimize the network connectivity for green and intelligent transportation system
Green Energy and Intelligent Transportation 2024, 3(3)
Published: 27 April 2024
Abstract Collect

LoRa technology contributes to green energy by enabling efficient, long-range communication for the Internet of Things (IoT). This paper addresses the challenges related to coverage range in outdoor monitoring systems utilizing LoRa, where the network performance is affected by the density of gateways (GWs) and end devices (EDs), as well as environmental conditions. To mitigate interference, data throughput losses, and high-power consumption, the proposed spreading factor (SF) and hybrid (data rate|SF) models dynamically adjust the transmission parameters. The orchestration of concurrent data modifications within the network server (NS) is crucial for uninterrupted communication between GWs and EDs, especially in monitoring electric vehicle (EV) stations to reduce traffic congestion and pollution. Employing K-means and density-based spatial clustering of applications with noise (DBSCAN) algorithms optimizes ED allocation, averts data congestion, and improves the signal-to-interference noise ratio (SINR). These methods ensure seamless information reception by meticulously allocated EDs across various GW combinations. To estimate the free-space losses (FSL), a log-distance path loss model (log-PL) is used. Exploring various bandwidths (BWs), bidirectional communications, and duty cycles (DCs) helps to prevent saturation, thus prolonging the operational lifespan of EDs. Empirical findings reveal a notable packet rejection rate (PRR) of 0% for the DBSCAN (hybrid model). In contrast, the K-means exhibits a PRR ranging from 5% (hybrid model) to 35.29% (SF model) for the ten GWs combination. Notably, the network saturation is reduced to 10.185% and 9.503%, respectively, highlighting an improvement in the average efficiency of slotted ALOHA (91.1%) and pure ALOHA (90.7%). These enhancements increase the lifespan of EDs to 15,465.27 days.

Total 2