Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
September 6, 2026Open Access

Adaptive AI Inference Optimization: A Comparative Simulation Study of Static, Reactive, Forecast-Based, and Optimization-Based Autoscaling

View Full Paper
Ask AI
Bookmark
Share

Authors

DCDivyashri Chinchole

Discussion

Loading...

Member takes

Overview

Simulation study reveals optimization-based autoscaling minimizes cost without service violations under dynamic AI workloads, highlighting efficiency gains over static provisioning.

Key Points

  • Evaluate and compare the performance of static, reactive, forecast-based, and optimization-based autoscaling strategies under dynamic AI inference workloads.
  • Simulated an AI inference infrastructure across 1,440 time steps using an identical synthetic workload.
  • Evaluated four autoscaling strategies: static provisioning, reactive autoscaling, forecast-based autoscaling, and optimization-based autoscaling.
  • Measured simulated infrastructure cost, estimated latency, SLA violations, active instance utilization, throughput, and queue behavior.
  • Optimization-based autoscaling achieved the lowest simulated total infrastructure cost of 41.10 while recording zero SLA violations and a maximum queue length of zero.
  • Reactive autoscaling achieved zero SLA violations but incurred a higher simulated infrastructure cost of 60.12.
  • Forecast-based autoscaling yielded a simulated cost of 45.65 with 5 SLA violations, whereas static provisioning produced the highest cost of 72.00 and 195 SLA violations.

Cite This Study

Divyashri Chinchole (2026) studied this question.

synapsesocial.com/papers/6a9d1ee128139818eab2209fhttps://doi.org/10.5281/zenodo.22292271
View Full Paper
Ask AI
Bookmark
Share