Key points are not available for this paper at this time.
Decentralized compute markets require autonomous agents to negotiate heterogeneous resources under budget constraints, stochastic supply, and strategic interaction. We present Agora-RL, a reproducibility-first benchmark for repeated negotiation of GPU, memory, and bandwidth through token-denominated double auctions. The study asks two questions: how standard MARL baselines rank when reward, social welfare, and inequality are evaluated jointly; and whether a transparent benchmark protocol can make such comparisons auditable. PPO, MAPPO, MADDPG, and IQL are evaluated with matched 300-episode training budgets, 30 deterministic evaluation episodes, and 12 random seeds. Using percentile-bootstrap 95% confidence intervals, MAPPO achieves the highest reward (0.0140 0.0124, 0.0154) and social welfare (0.0952 0.0854, 0.1045), whereas IQL yields the lowest Gini coefficient (0.4477 0.4360, 0.4613). Secondary diagnostics show that reward leadership does not imply fairness, equilibrium closeness, or communication robustness. The contribution is an empirical benchmark and audit protocol rather than a new auction theorem or blockchain settlement layer.
Gergov et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: