Randomized trial explores cold-start issues in AI applications, indicating potential solutions for scaling efficiency.
Serverless computing promises to run AI-powered, cloud-native applications with no idle cost and automatic, event-driven scaling, but the serverless model collides with the realities of modern AI workloads: large models, accelerator dependence, and multi-second cold starts. This paper presents a solution architecture for building AI applications on serverless primitives. We organize the system into four event-driven layers — application, orchestration, AI services, and data/state — in which loosely coupled functions and managed services are triggered by events and scale independently from zero. We analyze the central tension: cold-start latency, which for large models can reach tens of seconds to minutes, versus the scale-to-zero economics that make serverless attractive. We give a cost model that exposes the crossover beyond which reserved capacity is cheaper than pay-per-use, a cold-start latency model, and a survey of mitigation strategies (quantization, snapshotting/pre-loading, and provisioned concurrency) with their latency–cost trade-offs. Illustrative results, consistent with reported figures, show that an event-driven serverless design can serve bursty AI workloads at near-zero idle cost while bounding tail latency. Public sources are cited throughout.
No takes yet. Share an insight, caveat, or question.
Sushma Sunkollu Nagaraj (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: