Incremental multi-step Q-learning

202Citations
Citations of this article
73Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

This paper presents a novel incremental algorithm that combines Q-learning, a well-known dynamic-programming based reinforcement learning method, with the TD(λ) return estimation process, which is typically used in actor-critic learning, another well-known dynamic-programming based reinforcement learning method. The parameter λ is used to distribute credit throughout sequences of actions, leading to faster learning and also helping to alleviate the non-Markovtan effect of coarse state-space quantization. The resulting algorithm, Q(λ)-learning, thus combines some of the best features of the Q-learning and actor critic learning paradigms. The behavior of this algorithm has been demonstrated through computer simulations. © 1996 Kluwer Academic Publishers,.

Cite

CITATION STYLE

APA

Peng, J., & Williams, R. J. (1996). Incremental multi-step Q-learning. Machine Learning, 22(1–3), 283–290. https://doi.org/10.1007/BF00114731

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free