PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 15, 2022Journal of Systems Architecture19 citationsOpen Access

Reformulating the direct convolution for high-performance deep learning inference on ARM processors

View Full Paper
SBSergio BarrachinaACAdrián CastellóMDManuel F. Dolz

Key Points

Key points are not available for this paper at this time.

Abstract

We present two high-performance implementations of the convolution operator via the direct algorithm that outperform the so-called lowering approach based on the im2col transform plus the gemm kernel on an ARMv8-based processor. One of our methods presents the additional advantage of zero-memory overhead while the other employs an additional yet rather moderate workspace, substantially smaller than that required by the im2col+gemm solution. In contrast with a previous implementation of a similar zero-memory overhead direct convolution, this work exhibits the key advantage of preserving the conventional NHWC data layout for the input/output activations of the convolution layers.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Barrachina et al. (2022) studied this question.

synapsesocial.com/papers/6a110e036da19daf83169cdbhttps://doi.org/10.1016/j.sysarc.2022.102806
Ask AI
Helpful
Bookmark
Share
View Full Paper