Experimental analysis reveals Thompson sampling outperforms ETC and UCB in dynamic scenarios, suggesting adaptive mechanisms are crucial.
This paper comprehensively compares the performance of three multi-armed bandit (MAB) algorithms, Epsilon-Then-Commit (ETC), upper confidence bound (UCB), and Thompson sampling (TS), for video recommendation in dynamic environments. Using real TikTok interaction data, their performance is evaluated in both static and dynamic scenarios, where user preferences change. Experimental results show that Thompson sampling performs best, while ETC can only re-adapt to the environment after the environment changes due to its fixed exploration strategy. UCB shows moderate adaptability and relies on the adjustment of confidence intervals. TS's Bayesian approach can naturally balance exploration and exploitation without manual parameter tuning, and TS can achieve faster convergence recovery than the other two algorithms. Although TS has high computational overhead and long running time, its robustness in dynamic scenarios proves its application value in practical recommendation systems. This study highlights the importance of adaptive exploration mechanisms in dealing with non-stationary user behaviours and provides a reference for deploying multi-armed bandit algorithms on large-scale platforms. Future research directions include combining contextual information and hybrid models with deep learning techniques.
No takes yet. Share an insight, caveat, or question.
S. Li (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: