SNAP-UQ reduces flash and latency in TinyML by estimating uncertainty in a single pass, highlighting its efficiency.
We introduce SNAP-UQ, a single-pass, label-free uncertainty method for TinyML that estimates risk from depth-wise next-activation prediction: tiny int8 heads forecast the statistics of the next layer from a compressed view of the previous one, and a lightweight monotone mapper turns the resulting surprisal into an actionable score. The design requires no temporal buffers, auxiliary exits, or repeated forward passes, and adds only a few tens of kilobytes to MCU deployments. Across vision and audio backbones, SNAP-UQ consistently reduces flash and latency relative to early-exit and deep ensembles (typically ~40--60% smaller and ~25--35% faster), with competing methods of similar accuracy often exceeding memory limits. In corrupted streams it improves accuracy-drop detection by several AUPRC points and maintains strong failure detection (AUROC ≈0.9) in a single pass. Grounding uncertainty in layer-to-layer dynamics yields a practical, resource-efficient basis for on-device monitoring in TinyML.
No takes yet. Share an insight, caveat, or question.
Lamaakal et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: