We present experimental evidence that token selection in large language models can be performed wit hout computing the full transformer forward pass. By replacing softmax-based token selection with a retrieval-and-voting mechanism, we achieve 91-100% offline accuracy and a sustained 99% non-BOS to ken rate in standalone inference at 671 tokens per second on a consumer-grade Mac Studio. Our resul ts challenge the default assumption that inference quality is inseparable from matrix multiplicatio n.
文国 韩 (Fri,) studied this question.