Sort:
Open Access Issue
A Task-oriented Dialogue Policy Learning Method of Improved Discriminative Deep Dyna-Q
Journal of Guangdong University of Technology 2023, 40(4): 9-17
Published: 01 July 2023
Abstract PDF (1.5 MB) Collect
Downloads:1

As a pivotal part of the task-oriented dialogue system, dialogue policy can be trained by using the discriminative deep Dyna-Q framework. However, the framework uses vanilla deep Q-network method in the direct reinforcement learning phase and adopts MLPs as the basic network of world model, which limits the efficiency and stability of the dialogue policy learning. In this paper, we purpose an improved discriminative deep Dyna-Q method for task-oriented dialogue policy learning. In the improved direct RL phase, we first employ a NoisyNet to improve the exploration method, and then combine the dual-stream architecture of Dueling Network, Double-Q Network and n-step bootstrapping to optimize the calculation of the Q values. Moreover, we design a soft-attention-based model to replace the MLPs in the world model. The experimental results show that our proposed method achieves better results than other baseline models in terms of task success rate, average dialog turns and average reward. We further validate the effectiveness of proposed method by conducting both ablation and robustness analysis.

Open Access Issue
Efficient Temporal Modeling for RGB-T Tracking
Journal of Guangdong University of Technology 2026, 43(2): 21-29
Published: 03 June 2025
Abstract PDF (1 MB) Collect
Downloads:0

RGB-Thermal (RGB-T) tracking methods utilize the complementarity of visible light and thermal infrared images to improve the accuracy of target tracking in the scenarios of low light conditions and adverse weather. However, most existing studies focus only on image-level appearance matching, making them difficult to cope with challenges of target deformation and interference under complex environments. To address this problem, a tracking method based on efficient temporal modeling is proposed. Firstly, the temporal information is modeled, and the feature fusion module is improved to process temporal information. Then, a lightweight adapter is used for fine-tuning to improve the feature extraction module for thermal infrared images, enhancing the model's ability to extract features from different modal information, reducing the computational memory usage, and improving the training efficiency. Finally, a dynamic template update and selection method is proposed to fully explore and utilize temporal information, thereby improving the model's performance. ETMTrack achieves state-of-the-art performance on three public datasets, and performs excellently in dealing with challenges such as occlusion and similar appearances, demonstrating the effectiveness and robustness of the tracking algorithm based on temporal modeling.

Total 2