Randomized trial demonstrates performance limits of mixture-of-expert models on consumer hardware, indicating system constraints.
Does a large Mixture-of-Experts (MoE) language model have to fit in RAM to run at usable speed on consumer hardware? We built a streaming inference engine for Apple Silicon — it keeps a model's attention/norm/embedding tier resident and pulls MoE expert weights off the SSD on demand through a segmented-LRU cache — and held every design decision to a live/kill gate fixed before the experiment ran. GPT-OSS-120B (117B params, ~63 GB on disk) decodes coherently at 1.04 tok/s on an 18 GB machine, a model 3.5x larger than total RAM; Qwen3-30B-A3B reaches 7.08 tok/s with a 6 GB expert cache. We report the first multi-token expert-routing predictability measurements on modern decoder MoEs, finding overlap that does not decay across an 8-token horizon. The paper's more useful half is five pre-registered negative results: shared-base + low-rank expert decomposition does not transfer to fine-grained MoEs; speculative decoding increases per-token disk traffic for expert streaming; and three engine optimizations (batched fetch, persistent expert bank, native C++/Metal fetch primitive) each miss their speed gate because the bottleneck is a per-layer CPU-GPU round-trip intrinsic to data-dependent fetch. A lightweight expert predictor cannot dodge it — the signal is real and grows with scale (+9.66 points at 120B vs ~0 at 1-3B active) but far too weak to matter. The >=3 tok/s target is unreachable for streaming MoE inference on this hardware class. We release the engine, every measurement, and every negative result.
No takes yet. Share an insight, caveat, or question.
Muharrem Yurtsever (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: