AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (297.2 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

A study of value iteration and policy iteration for Markov decision processes in Deterministic systems

Haifeng Zheng( )Dan Wang
School of Economics, Jinan University, Guangzhou 510632, Guangdong, China
Show Author Information

Abstract

In the context of deterministic discrete-time control systems, we examined the implementation of value iteration (VI) and policy (PI) algorithms in Markov decision processes (MDPs) situated within Borel spaces. The deterministic nature of the system's transfer function plays a pivotal role, as the convergence criteria of these algorithms are deeply interconnected with the inherent characteristics of the probability function governing state transitions. For VI, convergence is contingent upon verifying that the cost difference function stabilizes to a constant k ensuring uniformity across iterations. In contrast, PI achieves convergence when the value function maintains consistent values over successive iterations. Finally, a detailed example demonstrates the conditions under which convergence of the algorithm is achieved, underscoring the practicality of these methods in deterministic settings.

CLC number: 60J05, 60J10

References

【1】
【1】
 
 
AIMS Mathematics
Pages 33818-33842

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Zheng H, Wang D. A study of value iteration and policy iteration for Markov decision processes in Deterministic systems. AIMS Mathematics, 2024, 9(12): 33818-33842. https://doi.org/10.3934/math.20241613

610

Views

1

Downloads

1

Crossref

0

Web of Science

1

Scopus

Received: 04 August 2024
Revised: 11 November 2024
Accepted: 19 November 2024
Published: 15 December 2024
Copyright © 2024 by AIMS Mathematics

This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 Unported License. To view a copy of this license, visit http://creativecommons.org/licenses/by-nc-sa/4.0/