This study examines the possibility of using modern neural network models that perform human recognition from photos at a local enterprise with up to 30 people. The participating models are FaceNet, FaceNet512, SFace, ArcFace, DeepFace, VGG-Face and Dlib from the DeepFace Framework. A feature of the study is the variation in the number of source photos in the database (1, 4, 8, 16 pieces per person), on the basis of which we can draw a conclusion about the accuracy and ease of use of a particular model in a real task. Accuracy is assessed by evaluating the F1 score at precision> 0.75. At the same time, threshold is assessed for each case and the performance of the networks under consideration. The work provides tables and graphs of the dependences of F1 score on the number of people, the number of photos per person, the neural network model used and metrics. As a result, a conclusion is made about the possibility of using FaceNet and FaceNet512 (30 people, 16 photos per person, cosine, F1 > 0.96, 90–100 ms per photo) and SFace (30 people, 16 photos per person, cosine, F1 > 0.8, 17 ms per photo). Moreover, both FaceNet models can be used for a database of 30 people, even if there is one photo per person (cosine, F1 > 0.94).
No takes yet. Share an insight, caveat, or question.
Berezin et al. (2024) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: