Abstract
In this paper, we present a new algorithm that integrates recent advances in solving continuous bandit problems with sample-based rollout methods for planning in Markov Decision Processes (MDPs). Our algorithm, Hierarchical Optimistic Optimization applied to Trees (HOOT) addresses planning in continuous-action MDPs. Empirical results are given that show that the performance of our algorithm meets or exceeds that of a similar discrete action planner by eliminating the problem of manual discretization of the action space.
Cite
CITATION STYLE
Mansley, C., Weinstein, A., & Littman, M. L. (2011). Sample-based planning for continuous action Markov decision processes. In ICAPS 2011 - Proceedings of the 21st International Conference on Automated Planning and Scheduling (pp. 335–338). https://doi.org/10.1609/icaps.v21i1.13484
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.