Abstract
A class of simple adaptive allocation rules is proposed for the problem (often called the "multi-armed bandit problem") of sampling xi,...,XN sequentially from k populations with densities belonging to an exponential family, in order to maximize the expected value of the sum SN = X1 + +XN. These allocation rules are based on certain upper confidence bounds, which are developed from boundary crossing theory, for the k population parameters. The rules are shown to be asymptotically optimal as N - oo from both Bayesian and frequentist points of view. Monte Carlo studies show that they also perform very well for moderate values of the horizon N.
Cite
CITATION STYLE
Lai, T. L. (2007). Adaptive Treatment Allocation and the Multi-Armed Bandit Problem. The Annals of Statistics, 15(3). https://doi.org/10.1214/aos/1176350495
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.