Motivated by practical applications, chiefly clinical trials, we study the regret achievable for stochastic bandits under the constraint that the employed policy must split trials into a small number of batches. We propose a simple policy, and show that a very small number of batches gives close to minimax optimal regret bounds. As a byproduct, we derive optimal policies with low switching cost for stochastic bandits.
No takes yet. Share an insight, caveat, or question.
Perchet et al. (2016) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: