Key points are not available for this paper at this time.
The paper explores convex optimisation in online distributed bandit scenarios across directed graphs that vary over time. The architecture of online optimisation with convex function can be seen as a structured and repeated process in which each online participant can obtain the loss function after making a decision, and then evaluate the performance of this algorithm by optimising the gap in total loss obtained by the decision maker after making a decision and the total loss caused by the best decision, that is, the regret upper bound. Our research focuses on scenarios where decision-makers can't directly obtain complete gradient information, while the interaction information is established on time-varying directed imbalanced networks, with their graph matrices being non-doubly stochastic. In response to these challenges, we propose two approximate gradient methods considering stochastic perturbations, combined with a weight-balancing technique, to develop two projection-free online optimisation algorithms. Specifically, by selecting appropriate step sizes, the algorithms can achieve consensus among the estimates and obtain sublinear regret under the objective function's strong convexity. Furthermore, we validate the effectiveness of our algorithms by means of numerical analyses.
Shang et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: