PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 15, 2022Journal of Systems Architecture19 citationsOpen Access

Reformulating the direct convolution for high-performance deep learning inference on ARM processors

View Full Paper
SBSergio BarrachinaUniversitat Jaume IACAdrián CastellóValve (United States)MDManuel F. DolzUniversitat Jaume I

Key Points

Key points are not available for this paper at this time.

Abstract

We present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Barrachina et al. (2022) studied this question.

synapsesocial.com/papers/6a110e036da19daf83169cdbhttps://doi.org/10.1016/j.sysarc.2022.102806
Ask AI
Helpful
Bookmark
Share
View Full Paper