An online policy gradient algorithm for Markov decision processes with continuous states and actions

Yao Ma; Tingting Zhao; Kohei Hatano; Masashi Sugiyama

Conference ProceedingsOPEN ACCESS

An online policy gradient algorithm for Markov decision processes with continuous states and actions

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2014) 8725 LNAI(PART 2) 354-369

DOI: 10.1007/978-3-662-44851-9_23

1Citations

3Readers

Abstract

We consider the learning problem under an online Markov decision process (MDP), which is aimed at learning the time-dependent decision-making policy of an agent that minimizes the regret - the difference from the best fixed policy. The difficulty of online MDP learning is that the reward function changes over time. In this paper, we show that a simple online policy gradient algorithm achieves regret O(√T) for T steps under a certain concavity assumption and O(logT) under a strong concavity assumption. To the best of our knowledge, this is the first work to give an online MDP algorithm that can handle continuous state, action, and parameter spaces with guarantee. We also illustrate the behavior of the online policy gradient method through experiments. © 2014 Springer-Verlag.

Author supplied keywords

Cite

CITATION STYLE

APA

Ma, Y., Zhao, T., Hatano, K., & Sugiyama, M. (2014). An online policy gradient algorithm for Markov decision processes with continuous states and actions. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 8725 LNAI, pp. 354–369). Springer Verlag. https://doi.org/10.1007/978-3-662-44851-9_23

An online policy gradient algorithm for Markov decision processes with continuous states and actions

Abstract

Author supplied keywords

Cite

Register to see more suggestions