Generative Inverse Deep Reinforcement Learning for Online Recommendation

28Citations
Citations of this article
37Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Deep reinforcement learning enables an agent to capture users' interest through dynamic interactions with the environment. It uses a reward function to learn user's interest and to control the learning process, attracting great interest in recommendation research. However, most reward functions are manually designed; they are either too unrealistic or imprecise to reflect the variety, dimensionality, and non-linearity of the recommendation problem. This impedes the agent from learning an optimal policy in highly dynamic online recommendation scenarios. To address the above issue, we propose a generative inverse reinforcement learning approach that avoids the need of defining an elaborative reward function. In particular, we model the recommendation problem as an automatic policy learning problem. We first generate policies based on observed users' preferences and then evaluate the learned policy by a measurement based on a discriminative actor-critic network. We conduct experiments on an online platform, VirtualTB, and demonstrate the feasibility and effectiveness of our proposed approach via comparisons with several state-of-the-art methods.

Cite

CITATION STYLE

APA

Chen, X., Yao, L., Sun, A., Wang, X., Xu, X., & Zhu, L. (2021). Generative Inverse Deep Reinforcement Learning for Online Recommendation. In International Conference on Information and Knowledge Management, Proceedings (pp. 201–210). Association for Computing Machinery. https://doi.org/10.1145/3459637.3482347

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free