Technical disclosure demonstrates streaming large Mixture-of-Experts models efficiently on desktop systems, suggesting practical advancements in model serving.
MoE-Direct is a Windows-based serving stack that runs Mixture-of-Experts (MoE) language models larger than system memory on an ordinary desktop, by keeping the full model resident on NVMe storage and streaming only the experts each token actually requests, through unbuffered direct reads into a model-aware, budget-bounded RAM slot cache. This technical disclosure documents the shipped v0.2 implementation and its measurements — including a matched reference-system pair against llama.cpp's mmap path on Qwen3.5-122B (5.59–5.69 decoded tokens/s, 2.32–2.34x over the mmap baseline), lossless byte-identity verification of repacked expert weights, deterministic replay, and first-token-to-sustained operation of a 1T-class model (Kimi K2.6, 447 GB file) on a 32 GB RAM / 16 GB VRAM machine — alongside frozen specifications and roadmap directions, each marked by its actual maturity. It is published to create a dated, publicly accessible record of these techniques so that they remain freely usable.
No takes yet. Share an insight, caveat, or question.
tmxkzm1925-max (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: