AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

A Task-oriented Dialogue Policy Learning Method of Improved Discriminative Deep Dyna-Q

School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China
Guangzhou Xuanyuan Research Institute Co., Ltd., Guangzhou 510000, China
Show Author Information

Abstract

As a pivotal part of the task-oriented dialogue system, dialogue policy can be trained by using the discriminative deep Dyna-Q framework. However, the framework uses vanilla deep Q-network method in the direct reinforcement learning phase and adopts MLPs as the basic network of world model, which limits the efficiency and stability of the dialogue policy learning. In this paper, we purpose an improved discriminative deep Dyna-Q method for task-oriented dialogue policy learning. In the improved direct RL phase, we first employ a NoisyNet to improve the exploration method, and then combine the dual-stream architecture of Dueling Network, Double-Q Network and n-step bootstrapping to optimize the calculation of the Q values. Moreover, we design a soft-attention-based model to replace the MLPs in the world model. The experimental results show that our proposed method achieves better results than other baseline models in terms of task success rate, average dialog turns and average reward. We further validate the effectiveness of proposed method by conducting both ablation and robustness analysis.

References

【1】
【1】
 
 
Journal of Guangdong University of Technology
Pages 9-17

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Dai B, Zeng B, Wei P-f, et al. A Task-oriented Dialogue Policy Learning Method of Improved Discriminative Deep Dyna-Q. Journal of Guangdong University of Technology, 2023, 40(4): 9-17. https://doi.org/10.12052/gdutxb.220122

407

Views

1

Downloads

0

Crossref

Received: 20 July 2022
Published: 01 July 2023
© 2023 Editorial Office of Journal of Guangdong University of Technology

This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/).