Multinomial logit bandit with linear utility functions

Mingdong Ou; Nan Li; Shenghuo Zhu; Rong Jin

Conference Proceedings

Multinomial logit bandit with linear utility functions

Ou M
Li N
Zhu S
et al.

IJCAI International Joint Conference on Artificial Intelligence (2018) 2018-July 2602-2608

DOI: 10.24963/ijcai.2018/361

9Citations

15Readers

Get full text

Abstract

Multinomial logit bandit is a sequential subset selection problem which arises in many applications. In each round, the player selects a K-cardinality subset from N candidate items, and receives a reward which is governed by a multinomial logit (MNL) choice model considering both item utility and substitution property among items. The player's objective is to dynamically learn the parameters of MNL model and maximize cumulative reward over a finite horizon T. This problem faces the exploration-exploitation dilemma, and the involved combinatorial nature makes it non-trivial. In recent years, there have developed some algorithms by exploiting specific characteristics of the MNL model, but all of them estimate the parameters of MNL model separately and incur a regret no better than ÕNT which is not preferred for large candidate set size N. In this paper, we consider the linear utility MNL choice model whose item utilities are represented as linear functions of ddimension item features, and propose an algorithm, titled LUMB, to exploit the underlying structure. It is proven that the proposed algorithm achieves ÕdKT regret which is free of candidate set size. Experiments show the superiority of the proposed algorithm.

Cite

CITATION STYLE

APA

Ou, M., Li, N., Zhu, S., & Jin, R. (2018). Multinomial logit bandit with linear utility functions. In IJCAI International Joint Conference on Artificial Intelligence (Vol. 2018-July, pp. 2602–2608). International Joint Conferences on Artificial Intelligence. https://doi.org/10.24963/ijcai.2018/361

Multinomial logit bandit with linear utility functions

Abstract

Cite

Register to see more suggestions