The multi-armed bandit problem models an agent that simultaneously attempts to acquire new knowledge (exploration) and optimize his decisions based on existing knowledge (exploitation). The agent attempts to balance these competing tasks in order to maximize his total value over the period of time considered. There are many practical applications of bandit algorithms, including clinical trials, adaptive routing or portfolio design. Over the last decade there has been an increased interest in developing bandit algorithms to address specific issues in recommender systems, such as improved product recommendation, the cold start problem, or personalization. The aim of this tutorial is to provide a brief introduction to the bandit problem with an overview of the various applications of bandit algorithms in recommendation.
No takes yet. Share an insight, caveat, or question.
Barraza‐Urbina et al. (2020) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: