Abstract
We propose an approach to solving Markov Decision Processes with non-Markovian rewards specified in Linear Temporal Logic interpreted over finite traces (LTLf). Our approach integrates automata representations of LTLf formulae into compiled MDPs that can be solved by off-the-shelf MDP planners, exploiting reward shaping to help guide search. Experiments with state-of-the-art UCT-based MDP planner PROST show automata-based reward shaping to be an effective method to guide search, producing solutions of superior quality, while maintaining policy optimality guarantees.
Cite
CITATION STYLE
Camacho, A., Chen, O., Sanner, S., & McIlraith, S. A. (2017). Non-markovian rewards expressed in LTL: Guiding search via reward shaping. In Proceedings of the 10th Annual Symposium on Combinatorial Search, SoCS 2017 (Vol. 2017-January, pp. 159–160). AAAI press. https://doi.org/10.1609/socs.v8i1.18421
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.