Randomized trial demonstrates enhanced inference efficiency in edge-cloud networks, indicating effectiveness of the approach.
Visual Generative Artificial Intelligence (GenAI) has emerged as a promising solution to deliver visually stunning content. To achieve seamless synergy between the generic and specialized GenAI models, edge-cloud collaborative net-works require an adaptive framework that integrates real-time perception, continual learning, and autonomous optimization. In this work, we develop an Agentic Deep Reinforcement Learning (DRL) framework, where a Large Language Model (LLM)-enabled agent perceives network states and allocates resources contextually. Next, we formulate a joint optimization problem of inference scheme selection and resource orchestration, aiming to minimize time-average inference latency subject to the inference task queue stability constraint. Based on Lyapunov optimization, we first transform the original long-term optimization problem into several deterministic sub-problems. Then, a DRL-based Inference Task Scheduling (DRL-ITS) algorithm is developed to solve the sub-problems in each time slot, where an LLM-enabled agent provides rich prior knowledge and accurate semantic interpretation for resource allocation. Finally, we provide theoretical and simulation evaluations to demonstrate that the DRL-ITS algorithm can obtain faster convergence and reduce inference latency by comparing it with other benchmark schemes.
No takes yet. Share an insight, caveat, or question.
Li et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: