Exploiting multi-step sample trajectories for approximate value iteration

Robert Wright; Steven Loscalzo; Philip Dexter; Lei Yu

Conference ProceedingsOPEN ACCESS

Exploiting multi-step sample trajectories for approximate value iteration

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2013) 8188 LNAI(PART 1) 113-128

DOI: 10.1007/978-3-642-40988-2_8

6Citations

9Readers

Abstract

Approximate value iteration methods for reinforcement learning (RL) generalize experience from limited samples across large state-action spaces. The function approximators used in such methods typically introduce errors in value estimation which can harm the quality of the learned value functions. We present a new batch-mode, off-policy, approximate value iteration algorithm called Trajectory Fitted Q-Iteration (TFQI). This approach uses the sequential relationship between samples within a trajectory, a set of samples gathered sequentially from the problem domain, to lessen the adverse influence of approximation errors while deriving long-term value. We provide a detailed description of the TFQI approach and an empirical study that analyzes the impact of our method on two well-known RL benchmarks. Our experiments demonstrate this approach has significant benefits including: better learned policy performance, improved convergence, and some decreased sensitivity to the choice of function approximation. © 2013 Springer-Verlag.

Cite

CITATION STYLE

APA

Wright, R., Loscalzo, S., Dexter, P., & Yu, L. (2013). Exploiting multi-step sample trajectories for approximate value iteration. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 8188 LNAI, pp. 113–128). https://doi.org/10.1007/978-3-642-40988-2_8

Exploiting multi-step sample trajectories for approximate value iteration

Abstract

Cite

Register to see more suggestions