Pre-registered benchmark reveals bounded 8-bit token caching cuts served latency up to 99.8% in generative models, indicating efficient edge inference.
A Decision-Oriented Token as a Bounded, Threshold-Free Cache Key for Context-Conditioned Generative Edge Outputs: In-the-Loop Validation and Comparison Against Vector-Quantization, Learned, and Semantic Caches Randolph James Ferlic, M.D. and Kimberly Kate Ferlic (Fieldstone Analytics, LLC, Austin, TX, USA) Preprint · Zenodo DOI: 10.5281/zenodo.22148612 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract Generative models are moving to the edge to serve transient, context-conditioned ("flash") applications, where much of the cost, energy, and latency is paid outside generation — on the input, invocation, reuse, and transmission margins. We study whether a single decision-oriented token can reclaim that budget as a reuse-and-anticipation layer, without modifying the generative model. A previously filed class-discriminant codebook encoder reduces each window of a multivariate context stream to one 8-bit token chosen to preserve the downstream decision; we characterize that token, on five real public datasets (UCI HAR, Air Quality, Appliances Energy, Metro Interstate Traffic, Beijing PM2.5) under pre-registered protocols, primarily as a cache key for generative outputs. Main results. (i) As a cache key the token is bounded (a ≤K-entry cache and ≤K generations by construction), threshold-free (exact-hash lookup, no similarity threshold to tune), and decision-oriented — it cache-hits on decision-equivalent states. At equal bits it beats a generic unsupervised vector-quantization key by +0.08–0.185 decision fidelity and a learned VQ-VAE (TOTEM-style) key by +0.10–0.12; against a semantic cache (embedding + tuned threshold) it dominates at the efficient low-generation operating point on learnable targets (matching fidelity at ~20–40× fewer generations with a bounded cache), while the semantic cache reaches a higher fidelity ceiling on one dataset. (ii) With real generative models in the loop (Qwen2.5-0.5B and the 3× larger 1.5B-Instruct), across two datasets and both classification and free-form-advisory outputs — eight cells, with the decision's target held out of the prompt — conditioning on the token's codebook prototype stays within 0.05 of full-window conditioning in all eight cells, and token-cached serving exactly matches fresh generation in all eight, validating compression-preserves-quality and caching-preserves-quality end-to-end (these small models are weak at the decision itself, and the token's denoising edge on the 0.5B model fades to parity at 1.5B). (iii) The token's transition structure supports speculative prefetch that hides ~28–57% of gated-regeneration latency (55%+ on three of five datasets) using a small order-1 model. Honestly reframed and reported verbatim. The token is not a bandwidth compressor: against a lossless coder its on-wire advantage is only ~1.2–2.4×; its compression value is being a fixed-1-byte decision-oriented key and a ~19–96× reduction in real LLM prefill tokens (GPT-4/GPT-4o tokenizer). Invocation gating confers no unique advantage — a crude per-channel change detector gates as well or better; the token's value is that one representation serves caching, gating, and prefetch. The fidelity edge requires a labeled calibration set; a persistent token memory does not improve point decisions, is a poor generation-budget allocator via novelty, and discards trend magnitude; caching fidelity is capped by target predictability. The contribution is a decision-oriented, bounded, threshold-free reuse layer for generative edge output — validated in the loop and honestly bounded. Highlights · Decision-oriented cache key — one 8-bit token keys generative-output reuse; at equal bits it beats a generic unsupervised vector-quantization key (+0.08–0.185 decision fidelity), a learned VQ-VAE (TOTEM-style) tokenizer (+0.10–0.12), and semantic caching at the efficient low-generation operating point. · Bounded and threshold-free — a ≤K-entry cache and ≤K generations by construction, exact-hash O(1) lookup with no similarity threshold to tune; a deterministic worst-case-cost guarantee a similarity cache cannot offer. · Premise validated in the loop — real generators (Qwen2.5-0.5B and the 3× larger 1.5B-Instruct) across two datasets and both classification and free-form-advisory outputs (eight cells, decision target held out): conditioning on the token prototype stays within 0.05 of full-window conditioning in all eight cells, and token-cached serving exactly matches fresh generation in all eight. · On-device measurement (Apple M5 Pro) — encoder 4–13 µs/window, cache lookup ~44 ns, generation 127 ms (0.5B) / 282 ms (1.5B); the encoder is ≈20,000× cheaper than a generation and gated + cached serving cuts served latency by a measured 98.8–99.8%, matching the invocation-count model to 0.00 percentage points. · Speculative prefetch — a kilobyte-scale order-1 token-transition model hides ~28–57% of gated-regeneration latency (55%+ on three of five datasets), with a time-shuffle placebo confirming real temporal structure. · Honest two-number compression — the token is not a bandwidth compressor (a lossless coder competes to within ~1.2–2.4× on-wire); its compression value is being a fixed-1-byte decision-oriented, cache-keyable index plus a ~19–96× reduction in real LLM prefill tokens (GPT-4/GPT-4o tokenizer). · Honest negatives and reframes (reported verbatim) — invocation gating is a free rider (a crude change detector gates as well or better); persistent token memory does not aid point decisions; novelty is a poor generation-budget allocator; the token discards trend magnitude; the discriminant fidelity edge requires labels; caching fidelity is predictability-capped; and the small in-loop models are themselves weak decision-makers on these numeric tasks. What this record contains · Manuscript_Paper41.pdf — the manuscript (with the five figures embedded). · PAPER_41_ZENODO_ARCHIVE.zip — the reproducibility archive: the nine frozen pre-registrations, the eleven per-phase runners (including the two-model generator-in-the-loop and the on-device latency/energy measurement, plus the compression/gating-baseline and learned-VQ-VAE runners), the two figure-rebuild scripts, the per-experiment result JSON files, and the five figures. All paths and identifiers are scrubbed and leak-scanned per the campaign deposit discipline; no raw benchmark data is redistributed (all datasets are public; download instructions are in the archive README). Cite as R. J. Ferlic and K. K. Ferlic, "A decision-oriented token as a bounded, threshold-free cache key for context-conditioned generative edge outputs," Zenodo, 2026, doi: 10.5281/zenodo.22148612. License and patent notice Released under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). Consistent with that license, no patent, patent-application, or other intellectual-property right of the authors is licensed, waived, granted, or otherwise conveyed by this deposit; the methods described — including the class-discriminant single-token codebook encoder, its multi-token product-quantization variant, its unsupervised nearest-centroid-distance monitoring mode, and the token-keyed output-caching and speculative-prefetch methods (the last the subject of U.S. Provisional Application No. 64/142,667, filed prior to this deposit) — are the subject of filed and pending U.S. patent applications. Licensing inquiries: randolphf@fieldstoneanalyticsllc.com. Companion deposits (spiral-domain-encoder-campaign) · Class-discriminant codebook construction for single-token signal compression: doi:10.5281/zenodo.20788187 · Single-token industrial sensor substrate: doi:10.5281/zenodo.20854722 · Deterministic multi-token token ladder: doi:10.5281/zenodo.22003179 · A universal, options-free, one-byte market-state primitive: doi:10.5281/zenodo.22116173 Keywords generative edge; on-device inference; class-discriminant codebook; output caching; cache key; vector quantization; VQ-VAE; semantic caching; speculative prefetch; dwell-time law; ephemeral applications; token economics; energy efficiency; edge AI; large-language-model caching; on-device latency; pre-registration; honest negatives
No takes yet. Share an insight, caveat, or question.
Ferlic et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: