Abstract Visual Generative Artificial Intelligence (GenAI) has emerged as a promising solution to deliver visually stunning content. To achieve seamless synergy between the generic and specialized GenAI models, edge-cloud collaborative net-works require an adaptive framework that integrates real-time perception, continual learning, and autonomous optimization. In this work, we develop an Agentic Deep Reinforcement Learning (DRL) framework, where a Large Language Model (LLM)-enabled agent perceives network states and allocates resources contextually. Next, we formulate a joint optimization problem of inference scheme selection and resource orchestration, aiming to minimize time-average inference latency subject to the inference task queue stability constraint. Based on Lyapunov optimization, we first transform the original long-term optimization problem into several deterministic sub-problems. Then, a DRL-based Inference Task Scheduling (DRL-ITS) algorithm is developed to solve the sub-problems in each time slot, where an LLM-enabled agent provides rich prior knowledge and accurate semantic interpretation for resource allocation. Finally, we provide theoretical and simulation evaluations to demonstrate that the DRL-ITS algorithm can obtain faster convergence and reduce inference latency by comparing it with other benchmark schemes.
Li et al. (Mon,) studied this question.