Deep reinforcement learning via past-success directed exploration

2Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

The balance between exploration and exploitation has always been a core challenge in reinforcement learning. This paper proposes “past-success exploration strategy combined with Softmax action selection”(PSE-Softmax) as an adaptive control method for taking advantage of the characteristics of the online learning process of the agent to adapt exploration parameters dynamically. The proposed strategy is tested on OpenAI Gym with discrete and continuous control tasks, and the experimental results show that PSE-Softmax strategy delivers better performance than deep reinforcement learning algorithms with basic exploration strategies.

Cite

CITATION STYLE

APA

Liu, X., Xu, Z., Cao, L., Chen, X., & Kang, K. (2019). Deep reinforcement learning via past-success directed exploration. In 33rd AAAI Conference on Artificial Intelligence, AAAI 2019, 31st Innovative Applications of Artificial Intelligence Conference, IAAI 2019 and the 9th AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019 (pp. 9979–9980). AAAI Press. https://doi.org/10.1609/aaai.v33i01.33019979

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free