Technical report details cost-effective AI inference setup in small labs using varied consumer machines.
This technical report describes Hayula Labs' fleet strategy for orchestrating AI inference across a heterogeneous collection of consumer-grade machines. We detail the architecture, operational metrics, and cost analysis of a production deployment spanning five machines in Kuwait (Mac Studio M2 Ultra 192GB, Acer Nitro V14 with RTX 4050 6GB, Framework Desktop, legacy workstations, NAS) with a unified routing layer. Over 90 days of operation, the fleet processed 3.9M inference requests with 99.4% uptime and a p95 latency of 8.2s. Cost analysis shows a fleet TCO of $5,055/year versus $187,200/year for equivalent cloud API throughput — a 97.3% savings. We compare our approach to Petals, FlexGen, and DeepSpeed-Inference, identifying task-typed allocation as the key differentiator for heterogeneous hardware. This report is intended as a practical reference for small labs and organizations building cost-effective inference infrastructure.
No takes yet. Share an insight, caveat, or question.
Yahya Saqban (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: