Online markov decision processes

195Citations
Citations of this article
79Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We consider a Markov decision process (MDP) setting in which the reward function is allowed to change after each time step (possibly in an adversarial manner), yet the dynamics rfixed. Similar to the experts setting, we address the question of how well an agent can do when compared to the reward achieved under the best stationary policy over time. We provide efficient algorithms, which have regret bounds with no dependence on the size of state space. Instead, these bounds depend only on a certain horizon time of the process and logarithmically on the number of actions. ©2009 INFORMS.

Cite

CITATION STYLE

APA

Even-Dar, E., Kakade, S. M., & Mansour, Y. (2009). Online markov decision processes. Mathematics of Operations Research, 34(3), 726–736. https://doi.org/10.1287/moor.1090.0396

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free