Value-difference based exploration: Adaptive control between epsilon-greedy and softmax

Michel Tokic; Günther Palm

Conference Proceedings

Value-difference based exploration: Adaptive control between epsilon-greedy and softmax

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2011) 7006 LNAI 335-346

DOI: 10.1007/978-3-642-24455-1_33

159Citations

161Readers

Get full text

Abstract

This paper proposes "Value-Difference Based Exploration combined with Softmax action selection" (VDBE-Softmax) as an adaptive exploration/exploitation policy for temporal-difference learning. The advantage of the proposed approach is that exploration actions are only selected in situations when the knowledge about the environment is uncertain, which is indicated by fluctuating values during learning. The method is evaluated in experiments having deterministic rewards and a mixture of both deterministic and stochastic rewards. The results show that a VDBE-Softmax policy can outperform ε-greedy, Softmax and VDBE policies in combination with on- and off-policy learning algorithms such as Q-learning and Sarsa. Furthermore, it is also shown that VDBE-Softmax is more reliable in case of value-function oscillations. © 2011 Springer-Verlag.

Cite

CITATION STYLE

APA

Tokic, M., & Palm, G. (2011). Value-difference based exploration: Adaptive control between epsilon-greedy and softmax. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 7006 LNAI, pp. 335–346). https://doi.org/10.1007/978-3-642-24455-1_33

Value-difference based exploration: Adaptive control between epsilon-greedy and softmax

Abstract

Cite

Register to see more suggestions