Computational evaluation demonstrates improved solution quality and search convergence in vehicle routing problems, indicating the value of bandit-driven cooperation.
This work focuses on the insertion and evaluation of a reinforcement learning algorithm based on the Multi-Armed Bandit method within a multi-agent framework for combinatorial optimization. The algorithm was integrated into the cooperation structure of the framework to manage the pool used for sharing solutions among agents and to reduce the occurrence of inefficient steps during the cooperation process, guiding agents toward the selection of better solutions. To enable and support this integration, the framework’s cooperation structure and the diversity of agents were also analyzed. The cooperation structure intermediates the communication among agents, while diversity is ensured by implementing agents with different behaviors in the framework. The experiments were conducted using the Vehicle Routing Problem with Time Windows as the test problem. The results of computational experiments showed that the cooperation structure of the framework and the diversity of agents are of paramount importance, enhancing the quality of the objective function values and runtime achieved so far in tests carried out with the framework. Additionally, the results showed better convergence in the search structure of agents when the Multi-Armed Bandit Pool was incorporated into their cooperation structure, achieving improved runtimes and, in certain cases, enhancing the objective function values as this method prevented agents from exploring inappropriate spaces during the search process.
No takes yet. Share an insight, caveat, or question.
Silva et al. (2026) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: