PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 15, 20250 citationsOpen Access

Generalized Linear Bandits: Almost Optimal Regret with One-Pass Update

View Full Paper
YZYujie ZhangSXShuhang XuPZPeng Zhao

Key Points

  • Achieves nearly optimal regret while maintaining a constant $ ext{O}(1)$ time complexity per round.
  • Utilizes a tight confidence set for the online mirror descent estimator to enhance statistical efficiency.
  • Addresses the trade-off between computational efficiency and optimal regret in existing generalized linear bandit methods.
  • Developed through novel analysis utilizing the concept of mix loss from online prediction.

Abstract

We study the generalized linear bandit (GLB) problem, a contextual multi-armed bandit framework that extends the classical linear model by incorporating a non-linear link function, thereby modeling a broad class of reward distributions such as Bernoulli and Poisson. While GLBs are widely applicable to real-world scenarios, their non-linear nature introduces significant challenges in achieving both computational and statistical efficiency. Existing methods typically trade off between two objectives, either incurring high per-round costs for optimal regret guarantees or compromising statistical efficiency to enable constant-time updates. In this paper, we propose a jointly efficient algorithm that attains a nearly optimal regret bound with O (1) time and space complexities per round. The core of our method is a tight confidence set for the online mirror descent (OMD) estimator, which is derived through a novel analysis that leverages the notion of mix loss from online prediction. The analysis shows that our OMD estimator, even with its one-pass updates, achieves statistical efficiency comparable to maximum likelihood estimation, thereby leading to a jointly efficient optimistic method.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68ef858cc6a308ba063553d3https://doi.org/10.48550/arxiv.2507.11847
Ask AI
Helpful
Bookmark
Share
View Full Paper