Thompson Sampling for Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit

5Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

We study the real-valued combinatorial pure exploration of the multi-armed bandit (R-CPE-MAB) problem. In R-CPE-MAB, a player is given d stochastic arms, and the reward of each arm s ∈ {1, . . ., d} follows an unknown distribution with mean µs. In each time step, a player pulls a single arm and observes its reward. The player’s goal is to identify the optimal action π∗ = arg max µπ from a finite-sized realπ∈A valued action set A ⊂ Rd with as few arm pulls as possible. Previous methods in the R-CPE-MAB require enumerating all of the feasible actions of the combinatorial optimization problem one is considering. In general, since the size of the action set grows exponentially large in d, this is almost practically impossible when d is large. We introduce an algorithm named the Generalized Thompson Sampling Explore (GenTS-Explore) algorithm, which is the first algorithm that can work even when the size of the action set is exponentially large in d. We also introduce a novel problem-dependent sample complexity lower bound of the R-CPE-MAB problem, and show that the GenTS-Explore algorithm achieves the optimal sample complexity up to a problem-dependent constant factor.

Cite

CITATION STYLE

APA

Nakamura, S., & Sugiyama, M. (2024). Thompson Sampling for Real-Valued Combinatorial Pure Exploration of Multi-Armed Bandit. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, pp. 14414–14421). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/aaai.v38i13.29355

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free