Paying Down the Price of the Bottleneck: Label-Efficient Foundation Codebooks and a Soft Readout for a Deterministic Edge Decision Token Randolph James Ferlic, M.D. and Kimberly Kate Ferlic (Fieldstone Analytics, LLC, Austin, TX, USA) Preprint · Zenodo DOI: 10.5281/zenodo.22866039 · CC-BY 4.0 · Community: spiral-domain-encoder-campaign Abstract A companion characterization established that a frozen, deterministic, class-discriminant single-token encoder (features → Fisher-discriminant ⊕ PCA subspace → k-means codebook of ≤ 256 cells → nearest-centroid decision) pays two specific bills for its byte-scale, auditable, sub-milliwatt bottleneck: it is label-hungry (a cost-setter, not a few-shot learner) and it carries a large-N accuracy tax (its hard, piecewise-constant readout is coarser than a continuous head). This paper asks, with the same pre-registration discipline, whether either bill can be paid down without changing the deterministic runtime or the ≤ 256-entry auditable table — and at what honest cost, across eight tasks (ECG, industrial bearing vibration, surgical-robot kinematics, wearable EMG, financial volatility, continuous glucose monitoring; binary and multiclass; group-disjoint splits with bootstrap confidence intervals). Two levers work. (1) A foundation codebook — centroids, subspace, and a pre-calibrated per-cell table frozen on a large labeled population and deployed to new subjects, refitting only the ≤ 256-entry table on N few-shot labels under an empirical-Bayes shrinkage rule — is near its population ceiling at N ≈ 0 where a from-scratch codebook is at chance (+0.33–0.36 AUROC at N = 8), monotone and never-worse, decision-shift-robust across device/site/era/age (gaps < 0.025), and a safe few-shot recalibrator that beats Platt scaling at small N; it replicates across all five sensor domains. It matches a nearest-class-mean/prototype head and tops logistic regression and k-nearest-neighbors at small N, but does not universally win (a linear model on linearly-separable features can beat it), so we position it as label-efficient and auditable. (2) A soft readout — averaging the ≤ 256-entry table over the m nearest cells (the filed distance-vote; the hard cell assignment, the emitted 8-bit token, and the static table all unchanged) — recovers 79–92% of the continuous linear ceiling and matches or exceeds it on 2 of 5 domains, knee at m ≈ 8, a weighting scheme that provably does not matter, composing with the foundation and generalizing to multiclass and every modality tested (+0.10 to +0.37 over the constant readout at deployable budgets, never worse at m = 8). We then discipline the claim on the metric a reviewer would attack: the recovery is a genuine discrimination + decision gain (balanced accuracy, macro-F1 — the ranking gain translates at 46% to > 100% across tasks), not a calibration gain — the soft readout de-calibrates because averaging cells pulls the score off each cell's empirical frequency. That bill is payable: a one-parameter temperature recalibration fit on the same N labels restores calibration to at or below the constant readout's native level while leaving the decision provably unchanged (soft-vote ECE 0.36 → 0.05 on glucose, 0.37 → 0.01 on bearing vibration, at N = 64). An honest ceiling check shows the linear reference was fair — gradient-boosted and neural models on the same features are few-shot-limited and do not beat it at deployable budgets, though a continuous model retains a ≤ ~0.12 AUROC residual at larger N on the hardest tasks. Four alternative levers are reported as pre-registered negatives (teacher distillation at population scale, upstream supervised codebook refinement, a per-cell piecewise-linear tuner, a multi-token "burst" ensemble). This is explicitly a characterization of previously-described, filed methods; it discloses no new algorithmic subject matter. Highlights · Two bills, paid down from opposite sides, runtime untouched. The token's label-hunger is paid down offline by a foundation codebook; its accuracy tax is paid down at runtime by a soft readout. Neither touches the deterministic nearest-centroid computation or the ≤ 256-entry auditable table. · Label-efficient and never-worse. A frozen foundation codebook + empirical-Bayes shrinkage table is at its population ceiling from the first labels (+0.33–0.36 AUROC at N = 8 vs a fresh codebook at chance), monotone, shift-robust, and a safer few-shot recalibrator than Platt — replicated across five sensor domains. Bounded honestly: it matches the strongest simple few-shot head and tops the rest, but is not a universal winner. · The accuracy tax is recoverable — from the readout, not construction. A distance-vote over the m nearest cells (filed) recovers 79–92% of the continuous linear ceiling (matches/exceeds it on 2/5 domains), knee at m ≈ 8, weighting scheme provably irrelevant, generalizing to multiclass and six modalities, never hurting at m = 8. Four construction-side alternatives fail — the tax is intrinsic to the hard collapse. · The right metric, and the honest cost. The recovery is a real discrimination + decision gain (balanced accuracy / macro-F1), not a calibration gain — the soft readout de-calibrates. Measuring calibration separately is the discipline; conflating it with AUROC is the error. · The recommendation, demonstrated not asserted. One-parameter temperature recalibration restores the soft readout's calibration to ≤ the constant readout's native level with the decision provably unchanged (ECE 0.36 → 0.05 on glucose at N = 64). Platt scaling overfits the small label set; temperature is the safe choice. · An honest ceiling. Strong nonlinear models on the same features are themselves few-shot-limited and do not beat the linear reference at deployable budgets — and at small N the frozen foundation soft readout matches or beats a freshly-trained strong model — but a continuous model keeps a ≤ ~0.12 AUROC residual at larger N on the hardest tasks. "Recovers the tax" means recovers the linear-ceiling gap. What this record contains · Manuscript_Paper46.pdf — the manuscript with five figures embedded (the two-lever scorecard, the five-domain label-efficiency comparison, the accuracy↔auditability curve, the soft-vote advantage vs label-budget surface, and the post-hoc recalibration panel), and Manuscript_Paper46.docx, the editable source. · PAPER_46_ZENODO_ARCHIVE.zip — the reproducibility archive: the nine frozen pre-registrations, the experiment runners (the soft-vote m-sweep, the label-budget × cell-count surface, the metric-hygiene and reviewer-proofing rounds, the upstream-refinement / local-tuner / multi-token-ensemble negatives, the multi-domain foundation and reviewer-defense batteries, and the Modal-scale foundation / distillation / shrinkage / shift / recalibration runners), the shared feature-extractor loader, the figure builder, the seventeen per-experiment result records, the five figures, the manuscript source, and a README. All datasets are public; no raw benchmark data is redistributed (sources and a `PATH_TO_DATA` convention are in the README). All paths and identifiers are scrubbed and leak-scanned per the campaign deposit discipline. Cite as R. J. Ferlic and K. K. Ferlic, "Paying down the price of the bottleneck: label-efficient foundation codebooks and a soft readout for a deterministic edge decision token," Zenodo, 2026, doi: 10.5281/zenodo.22866039. License and patent notice Released under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). Consistent with that license, no patent, patent application, or other intellectual-property right of the authors is licensed, waived, granted, or otherwise conveyed by this deposit. This work characterizes previously-described methods and discloses no new algorithmic subject matter; distance-weighted voting, temperature and Platt scaling, prototype / nearest-class-mean classification, and empirical-Bayes shrinkage are established prior art, used only as tools. The methods characterized — the class-discriminant single-token codebook encoder and its nearest-centroid monitor, the multi-token / token-ladder / soft-readout mechanisms (the soft readout), and the supervised codebook-refinement mechanism (evaluated as a negative result) — are the subject of filed and pending U.S. patent applications held by the authors, including U.S. Provisional Application No. 64/095,354 (the encoder) and the multi-token / token-ladder readout and supervised-refinement applications (priority U.S. Application No. 19/467,303 and its continuations). The foundation-codebook-with-shrinkage edge deployment is published as a characterization of the already-filed codebook method — a prior-art review found it anticipated by the prototype / nearest-class-mean and training-free-calibration literatures — and no new subject matter is claimed for it. Per-deployment productization and deployment-selection know-how are not disclosed and are retained as trade secrets. © 2026 Fieldstone Analytics, LLC and the authors; all rights not expressly granted under CC-BY 4.0 are reserved. Licensing and collaboration inquiries: randolphf@fieldstoneanalyticsllc.com. Companion deposits (spiral-domain-encoder-campaign) · Class-discriminant codebook construction for single-token signal compression: doi:10.5281/zenodo.20788187 · Deterministic multi-token token ladder for channel-partition compression (the multi-token / soft-readout family): doi:10.5281/zenodo.22003179 · Label-free inference-time channel fusion for drift-robust single-token decisions: doi:10.5281/zenodo.22046713 · A decision-oriented token as a bounded, threshold-free cache key for generative edge outputs: doi:10.5281/zenodo.22148612 · The predictive reach of a decision token (forecasting/anticipation/fusion): doi:10.5281/zenodo.22736921 · Non-invertible but not anonymous: a privacy characterizatio
No takes yet. Share an insight, caveat, or question.
Ferlic et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: