Reinforcement Learning for POMDP Environments Using State Representation with Reservoir Computing

4Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

One of the challenges in reinforcement learning is regarding the partially observable Markov decision process (POMDP). In this case, an agent cannot observe the true state of the environment and perceive different states to be the same. Our proposed method uses the agent’s time-series information to deal with this imperfect perception problem. In particular, the proposed method uses reservoir computing to transform the time-series of observation information into a nonlinear state. A typical model of reservoir computing, the echo state network (ESN), transforms raw observations into reservoir states. The proposed method is named dual ESNs reinforcement learning, which uses two ESNs specialized for observation and action information. The experimental results show the effectiveness of the proposed method in environments where imperfect perception problems occur.

Cite

CITATION STYLE

APA

Yamashita, K., & Hamagami, T. (2022). Reinforcement Learning for POMDP Environments Using State Representation with Reservoir Computing. Journal of Advanced Computational Intelligence and Intelligent Informatics, 26(4), 562–569. https://doi.org/10.20965/jaciii.2022.p0562

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free