Fault-injection study demonstrates complete availability in a generative AI tool during provider outages, indicating the resilience of parallel racing and local fallbacks.
Applications built on third-party Generative AI APIs inherit the unreliability of those APIs. Providers rate-limit requests, exhaust free-tier quotas, deprecate models, and return server errors, any of which can break a dependent application. This paper presents an availability-oriented architecture that does not depend on any single provider. Instead of the common sequential-fallback approach, where a backup provider is tried only after the primary times out, the system queries several providers in parallel and returns the first successful response. It combines this parallel racing with per-provider health tracking, a circuit breaker that removes repeatedly failing providers, an in-memory cache for repeated inputs, and a deterministic local engine that generates a valid response with no external call when every provider fails. We implement the architecture as a Git commit-message generator, called IntelliCommit, and evaluate its behaviour under controlled fault injection, using a mock provider with a configurable failure rate and a fixed random seed so that results are reproducible on any machine. Across injected provider-failure rates of 30%, 60%, and 100% (50 requests each), every request returned a usable commit message, because retries and the deterministic local fallback ensure a response even when external provider attempts fail. Mean latency rises from approximately 1.1 to 2.9 seconds as the failure rate increases and more requests take the retry-then-fallback path. We report availability, defined as the share of requests returning a usable message, rather than message quality, which we do not evaluate in this work. All measurements, including per-request logs, are reproducible from the open-source repository. Index Terms—fault tolerance, large language model APIs, multi-provider architecture, availability, fault injection, graceful degradation.
No takes yet. Share an insight, caveat, or question.
Prashant Kumar Maurya (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: