Abstract
In this study, we extend the framework of semiparametric statistical inference introduced recently to reinforcement learning [1] to online learning procedures for policy evaluation. This generalization enables us to investigate statistical properties of value function estimators both by batch and online procedures in a unified way in terms of estimating functions. Furthermore, we propose a novel online learning algorithm with optimal estimating functions which achieve the minimum estimation error. Our theoretical developments are confirmed using a simple chain walk problem. © 2009 Springer Berlin Heidelberg.
Cite
CITATION STYLE
Ueno, T., Maeda, S. I., Kawanabe, M., & Ishii, S. (2009). Optimal online learning procedures for model-free policy evaluation. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 5782 LNAI, pp. 473–488). Springer Verlag. https://doi.org/10.1007/978-3-642-04174-7_31
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.