Abstract
This work proposes Dynamic Linear Epsilon-Greedy, a novel contextual multi-armed bandit algorithm that can adaptively assign personalized content to users while enabling unbiased statistical analysis.Traditional A/B testing and reinforcement learning approaches have trade-offs between empirical investigation and maximal impact on users.Our algorithm seeks to balance these objectives, allowing platforms to personalize content effectively while still gathering valuable data.Dynamic Linear Epsilon-Greedy was evaluated via simulation and an empirical study in the ASSISTments online learning platform.In simulation, Dynamic Linear Epsilon-Greedy performed comparably to existing algorithms and in ASSISTments, slightly increased students’ learning compared to A/B testing.Data collected from its recommendations allowed for the identification of qualitative interactions, which showed high and low knowledge students benefited from different content.Dynamic Linear Epsilon-Greedy holds promise as a method to balance personalization with unbiased statistical analysis.All the data collected during the simulation and empirical study are publicly available at https://osf.io/zuwf7/.
Author supplied keywords
Cite
CITATION STYLE
Prihar, E., Sales, A., & Heffernan, N. (2023). A Bandit You Can Trust. In UMAP 2023 - Proceedings of the 31st ACM Conference on User Modeling, Adaptation and Personalization (pp. 106–115). Association for Computing Machinery, Inc. https://doi.org/10.1145/3565472.3592955
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.