Accelerating TinyML Inference on Microcontrollers Through Approximate Kernels

Key Points

Key points are not available for this paper at this time.

Abstract

Despite the widespread adoption of energy-efficient microcontroller units (MCUs) in the Tiny Machine Learning (TinyML) domain, they face significant limitations in terms of performance and memory (RAM, Flash), especially when considering deep networks for complex classification tasks. In this work, we combine significance-aware computation skipping and software kernel design to accelerate the inference of approximate CNN models on MCUs. Our evaluation on an STM32-Nucleo board and 2 popular CNNs trained on the CIFAR-10 dataset shows that, compared to state-of-the-art exact inference, our Pareto optimal solutions can feature on average 21% latency reduction with no degradation in Top-1 classification accuracy, while for lower accuracy requirements, the corresponding reduction becomes even more pronounced.

Mark Helpful

Bookmark

Relay

Mark Helpful

Bookmark

Relay

Accelerating TinyML Inference on Microcontrollers Through Approximate Kernels

Key Points

Abstract

Cite This Study