Key points are not available for this paper at this time.
The paper is devoted to the study of the neural networks inference acceleration using the weights quantization and Intel OpenVINO Toolkit. At the same time, the study considers block architecture convolutional networks trained from scratch. In addition, it is shown that transfer learning makes it possible to obtain higher model accuracy. And their implementation on OpenVINO can significantly increase the processing performance. It is also shown that the use of OpenVINO provides significant acceleration without loss of performance for such networks, while quantization leads to significant loss of quality.
Andriyanov et al. (Mon,) studied this question.