In the context of deterministic discrete-time control systems, we examined the implementation of value iteration (VI) and policy (PI) algorithms in Markov decision processes (MDPs) situated within Borel spaces. The deterministic nature of the system's transfer function plays a pivotal role, as the convergence criteria of these algorithms are deeply interconnected with the inherent characteristics of the probability function governing state transitions. For VI, convergence is contingent upon verifying that the cost difference function stabilizes to a constant
Publications
- Article type
- Year
Article type
Year
Open Access
Research Article
Issue
AIMS Mathematics 2024, 9(12): 33818-33842
Published: 15 December 2024
Downloads:1
Total 1
京公网安备11010802044758号