Adaptive Treatment Allocation and the Multi-Armed Bandit Problem

  • Lai T
N/ACitations
Citations of this article
60Readers
Mendeley users who have this article in their library.

Abstract

A class of simple adaptive allocation rules is proposed for the problem (often called the "multi-armed bandit problem") of sampling xi,...,XN sequentially from k populations with densities belonging to an exponential family, in order to maximize the expected value of the sum SN = X1 + +XN. These allocation rules are based on certain upper confidence bounds, which are developed from boundary crossing theory, for the k population parameters. The rules are shown to be asymptotically optimal as N - oo from both Bayesian and frequentist points of view. Monte Carlo studies show that they also perform very well for moderate values of the horizon N.

Cite

CITATION STYLE

APA

Lai, T. L. (2007). Adaptive Treatment Allocation and the Multi-Armed Bandit Problem. The Annals of Statistics, 15(3). https://doi.org/10.1214/aos/1176350495

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free