Empirical benchmarking reveals a 120B MoE model can achieve 100 tok/s on sub-500 GB/s memory systems, indicating the feasibility of low-cost consumer AI appliances.
Consumer access to large language models is gated by the price of high-bandwidth GPU memory. I investigate whether a purpose-built appliance — a "player" holding a modest fast-memory tier and an accelerator, plus a cheap flashable NAND "disk" holding a mastered model image — can serve a ~120B-parameter mixture-of-experts model at 100 aggregate tokens/s across a 512K-token aggregate context budget. Measurements on Qwen3.6-35B-A3B (256 experts top-8 plus one shared, 3:1 linear-to-full hybrid attention, 1-layer MTP head) as an architectural proxy show: (1) expert routing is measurably concentrated (mean decode entropy 7.44 vs 8.00 uniform bits) and transition-predictable (+17–20 points over a static-popularity baseline), with an LRU expert cache reaching 93.7% hits at 75% residency (~22 MB of flash traffic per output token); (2) multi-token-prediction speculative acceptance is 90/82/66/45% at draft widths 1/2/4/8 and is invariant to context length, concurrency, GPU generation, and bf16-vs-FP8 weight format; (3) a same-format GPU ladder spanning 3.49× memory bandwidth shows per-request decode scaling at 90% of the bandwidth ratio in low-concurrency long-context modes. Calibrating an analytical bandwidth model with these measurements, the target contract fits within 350–500 GB/s of fast-memory bandwidth and 48–64 GB of capacity — LPDDR-class, not HBM-class, hardware. Total measurement cost was $8.01 of rented GPU time. All harnesses, routing traces, and raw results are bundled for reproduction.
No takes yet. Share an insight, caveat, or question.
Pranab Sarkar (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: