Motivated by cognitive radios, there has been a recent increase of interest in stochastic multi-player multi-armed bandits. In this context of cognitive radio's, autonomous players concurrently engage in arm or channel pulls, individually opti-mizing rewards. Complexity amplifies with potential collisions, wherein multiple players simultaneously select a common arm, resulting in zero collective reward. Our work centers on the Multiplayer Multi-Armed Bandit (MMAB) problem, involving M decision makers collaborating to maximize cumulative reward in cognitive radio application. Collision prompts players to adapt. We introduce RobustMMAB, a decentralized algorithm aiming to achieve regret akin to an optimal centralized algorithm while increasing resilience against selfish nodes.
No takes yet. Share an insight, caveat, or question.
Singh et al. (2024) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: