Non-markovian rewards expressed in LTL: Guiding search via reward shaping

40Citations
Citations of this article
18Readers
Mendeley users who have this article in their library.

Abstract

We propose an approach to solving Markov Decision Processes with non-Markovian rewards specified in Linear Temporal Logic interpreted over finite traces (LTLf). Our approach integrates automata representations of LTLf formulae into compiled MDPs that can be solved by off-the-shelf MDP planners, exploiting reward shaping to help guide search. Experiments with state-of-the-art UCT-based MDP planner PROST show automata-based reward shaping to be an effective method to guide search, producing solutions of superior quality, while maintaining policy optimality guarantees.

Cite

CITATION STYLE

APA

Camacho, A., Chen, O., Sanner, S., & McIlraith, S. A. (2017). Non-markovian rewards expressed in LTL: Guiding search via reward shaping. In Proceedings of the 10th Annual Symposium on Combinatorial Search, SoCS 2017 (Vol. 2017-January, pp. 159–160). AAAI press. https://doi.org/10.1609/socs.v8i1.18421

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free