PAC bounds for discounted MDPs

Tor Lattimore; Marcus Hutter

Conference Proceedings

PAC bounds for discounted MDPs

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2012) 7568 LNAI 320-334

DOI: 10.1007/978-3-642-34106-9_26

68Citations

41Readers

Get full text

Abstract

We study upper and lower bounds on the sample-complexity of learning near-optimal behaviour in finite-state discounted Markov Decision Processes (mdps). We prove a new bound for a modified version of Upper Confidence Reinforcement Learning (ucrl) with only cubic dependence on the horizon. The bound is unimprovable in all parameters except the size of the state/action space, where it depends linearly on the number of non-zero transition probabilities. The lower bound strengthens previous work by being both more general (it applies to all policies) and tighter. The upper and lower bounds match up to logarithmic factors provided the transition matrix is not too dense. © 2012 Springer-Verlag.

Author supplied keywords

Cite

CITATION STYLE

APA

Lattimore, T., & Hutter, M. (2012). PAC bounds for discounted MDPs. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 7568 LNAI, pp. 320–334). https://doi.org/10.1007/978-3-642-34106-9_26

PAC bounds for discounted MDPs

Abstract

Author supplied keywords

Cite

Register to see more suggestions