PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 10, 20240 citationsOpen Access

Fast UCB-type algorithms for stochastic bandits with heavy and super heavy symmetric noise

View Full Paper
YDYuriy DornAKAlexandr KatrutsaILIlgam Latypov

Key Points

Key points are not available for this paper at this time.

Abstract

In this study, we propose a new method for constructing UCB-type algorithms for stochastic multi-armed bandits based on general convex optimization methods with an inexact oracle. We derive the regret bounds corresponding to the convergence rates of the optimization methods. We propose a new algorithm Clipped-SGD-UCB and show, both theoretically and empirically, that in the case of symmetric noise in the reward, we can achieve an O (TKT T) regret bound instead of O (T^1{1+} K^{1+}) for the case when the reward distribution satisfies Eₗ ₃|X|^1+ ^1+ ( (0, 1]), i. e. perform better than it is assumed by the general lower bound for bandits with heavy-tails. Moreover, the same bound holds even when the reward distribution does not have the expectation, that is, when <0.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dorn et al. (2024) studied this question.

synapsesocial.com/papers/68e79adbb6db64358770b4cehttps://doi.org/10.48550/arxiv.2402.07062
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Data-Driven Upper Confidence Bounds with Near-Optimal Regret for Heavy-Tailed Bandits2024
  2. 2Beyond Primal-Dual Methods in Bandits with Stochastic and Adversarial Constraints2024
  3. 3Stochastic Bandits Robust to Adversarial Attacks2024
  4. 4Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits2025
  5. 5Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update2025