Finite Sample Analysis of Minmax Variant of Offline Reinforcement Learning for General MDPs

2Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

In this work, we analyze the finite sample complexity bounds for offline reinforcement learning with general state, general function space and state-dependent action sets. The algorithm analyzed does not require the knowledge of the data-collection policy as compared to earlier works. We show that one can compute an ϵ-optimal Q function (state-action value function) using O(1/ϵ4) i.i.d. samples of state-action-reward-next state tuples.

Cite

CITATION STYLE

APA

Regatti, J. R., & Gupta, A. (2022). Finite Sample Analysis of Minmax Variant of Offline Reinforcement Learning for General MDPs. IEEE Open Journal of Control Systems, 1, 152–163. https://doi.org/10.1109/OJCSYS.2022.3198660

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free