Abstract
In this work, we analyze the finite sample complexity bounds for offline reinforcement learning with general state, general function space and state-dependent action sets. The algorithm analyzed does not require the knowledge of the data-collection policy as compared to earlier works. We show that one can compute an ϵ-optimal Q function (state-action value function) using O(1/ϵ4) i.i.d. samples of state-action-reward-next state tuples.
Author supplied keywords
Cite
CITATION STYLE
Regatti, J. R., & Gupta, A. (2022). Finite Sample Analysis of Minmax Variant of Offline Reinforcement Learning for General MDPs. IEEE Open Journal of Control Systems, 1, 152–163. https://doi.org/10.1109/OJCSYS.2022.3198660
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.