AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (5.1 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Adversarial strategy generation integrating expert policies and multi-chain-of-thought reasoning

Yanan Ni1Runnan Qi1Kuihua Huang1( )Zongyuan Li2Xingxing Liang1
Laboratory for Big Data and Decision, National University of Defense Technology, Changsha 410073, China
College of Artificial Intelligence, Nankai University, Tianjin 300350, China
Show Author Information

Abstract

Objective

In adversarial environments, especially within highly dynamic real-time strategy (RTS) game scenarios, AI agents powered by large language models (LLMs) still face significant challenges in executing fine-grained tactical decision-making. Although LLMs exhibit strong language-based reasoning capabilities, they often lack real-time responsiveness and contextual adaptability when dealing with unit-level operations such as movement planning and attack prioritization. This study focuses on enhancing the strategic generation capabilities of LLM-driven agents in such complex tasks by integrating expert policies with multi-chain-of-thought (MCoT) reasoning. The aim is to strengthen their situational awareness and strategic coherence, thereby improving the accuracy, interpretability, and flexibility of decision-making.

Methods

This study proposes a method called AE-MCoT (adversarial strategy generation integrating expert policies and multi-chain-of-thought reasoning), which integrates multimodal inputs, expert policies, and MCoT reasoning to enhance the fine-grained decision-making capabilities of large language models (LLMs) in highly dynamic adversarial environments. The core of AE-MCoT lies in combining temporally segmented expert strategies with a parallel reasoning framework that explores multiple tactical paths. The expert policy module injects tactical priors into the LLM via natural language prompts embedded in the system prompt, enabling phase-focused reasoning throughout the game. Based on prior experiments and typical game rhythms, the gameplay is divided into three phases: early (0-6s), mid (6-20s), and late (>20s), corresponding respectively to goals such as evading early enemy contact, balancing offense and defense, and preserving unit survival at low health. These phase-specific strategies are marked with EARLY, MID, and LATE labels and dynamically invoked based on the in-game timestamp. Guided by these prompts, the model generates three distinct reasoning paths: aggressive (attack-first), conservative (survival-first), and balanced (offense-defense tradeoff), based on current environmental observations. A self-assessment mechanism is then employed to evaluate the effectiveness of each chain and select the optimal strategy for generating specific move-and-attack actions. Without the need for additional training, AE-MCoT enables high-precision, interpretable, and temporally coherent strategy generation, making it particularly suitable for fast-paced and complex adversarial scenarios.

Results

In the high-difficulty custom StarCraft Ⅱ scenario “1 Colossus vs. 32 Zerglings,” the AE-MCoT method demonstrated strong effectiveness. The full approach integrating expert policies and MCoT reasoning achieved a 95% win rate across 20 matches, surpassing both the single-chain and expert-ablated variants (each at 55%) and far outperforming the no-chain setting (5%). It reached a high kill-to-loss ratio (629), showing clear advantages in unit survivability and enemy suppression. Average kills per game reached 31.45, with 0.32 health ratio remaining upon victory, highlighting the model’s balance of offense and caution. In contrast, the chain-ablated version performed worst, with only 17.4 kills on average, the shortest match duration, and near-zero health retention. Under different enemy scales, AE-MCoT maintained robustness, achieving 65% wins against 40 enemies and 100% against 24, while the baseline agent failed in all cases. In an optimal match, the agent adjusted dynamically through early, mid, and late phases—initially evading encirclement, maintaining steady offense in midgame, and executing high-ground kiting for full elimination. The composite performance score (0.782) further confirms the method’s superior efficiency, survivability, and pacing in high-intensity fine-grained adversarial scenarios.

Conclusions

The results demonstrate the effectiveness of combining expert policies with MCoT reasoning for adversarial strategy generation. The proposed method improves the agent's ability to generate adaptive, fine-grained control strategies without additional training, making it a valuable tool for complex decision-making tasks. The integration of expert policies with MCoT reasoning ensures high adaptability and precision, enabling the model to adjust its strategy based on dynamic game conditions. This approach shows significant promise not only in RTS games but also in other complex, adversarial environments such as autonomous systems and military applications. Moreover, it highlights the potential for improving LLM-based agents' performance in highly dynamic adversarial environments without the need for retraining or extensive datasets.

CLC number: TP18 Document code: A Article ID: 1001-2486(2026)04-200-11

References

【1】
【1】
 
 
Journal of National University of Defense Technology
Pages 200-210

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ni Y, Qi R, Huang K, et al. Adversarial strategy generation integrating expert policies and multi-chain-of-thought reasoning. Journal of National University of Defense Technology, 2026, 48(4): 200-210. https://doi.org/10.11887/j.issn.1001-2486.25040020

4

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 14 April 2025
Published: 01 August 2026
© 2026 Journal of National University of Defense Technology

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).