Computational evaluation demonstrates token-free byte-level state space modeling achieves linear complexity, highlighting scalable alternatives to subword transformer architectures.
We present Stream, a token-free, position-free language model built entirely on state space model (SSM) recurrence operating directly on raw bytes. Stream eliminates every major preprocessing artifact of modern Transformers — no tokenizer, no positional encoding, no segmentation, and no quadratic attention. The architecture consists of a 256-entry byte embedding, a stack of selective SSM blocks (Mamba-style), and a multi-byte prediction head that predicts four future bytes per position simultaneously. Despite using only 256 vocabulary entries compared to the 50k+ of typical subword models, Stream achieves a validation loss of 1.69 at 4.43M parameters, compared to 1.20 for a similarly-sized Transformer (nanoGPT 8L/192D at 3.59M). The 0.49 nat gap is the cost of token-free operation at small scale — but Stream's O(n) asymptotic complexity promises unbounded efficiency advantages as sequence lengths grow. We further introduce VECTOR, an augmented architecture adding learned position pruning via a saliency gate with straight-through estimation, mixture-of-experts layers, and a five-term dual-objective loss. Finally, we present MoE-Stream, which augments Stream with sparse mixture-of-experts layers with gradient-aware expert lifecycle management.
No takes yet. Share an insight, caveat, or question.
Nishant Paudel (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: