Randomized trial demonstrates dynamic specialization in multi-domain serving on consumer hardware, suggesting improved efficiency.
Mixture-of-Experts models achieve efficient scaling by routing tokens to specialized sub-networks. However, traditional MoE requires training from scratch or expensive full-model fine-tuning. We propose Dense AdapterMoE: a hybrid architecture combining a frozen 7-8B base language model with 10-20 dynamically routed LoRA expert adapters, all loaded simultaneously on consumer hardware (Apple M2 Ultra, 192GB unified memory). Each expert adapter occupies only 37MB (2 million parameters at rank 8), and 20 experts total 740MB — representing 0.02% additional memory overhead over the base model. The system uses a learned routing mechanism to dispatch tokens to the most relevant expert, achieving dynamic specialization without the computational overhead of traditional MoE (2% FLOPs increase vs 100%+ for Mixtral-style MoE). Total memory footprint: approximately 17GB for the full system, leaving 175GB for a 1M+ token KV cache. This architecture enables 10-20 domain specializations within a single inference run, at 20-50 tok/s on M2 Ultra. We compare against Switch Transformer, Mixtral 8x7B, and dense fine-tuning approaches, identifying Dense AdapterMoE as optimal for multi-domain serving on memory-rich consumer hardware.
No takes yet. Share an insight, caveat, or question.
Yahya Saqban (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: