Randomized trial compares deep learning models for species identification, indicating efficient data processing methods for Lepidoptera.
The number of image‐based species records has increased exponentially in recent years. Deep learning models offer opportunities for efficient data processing, especially for species identification. Evaluating models on comprehensive datasets can provide insights into their suitability for species identification. Based on a high‐quality dataset with more than 500,000 images, we compared 40 models from 12 classes on the classification of 162 butterfly and moth species, and further analysed the results of the best model. The dataset covers most Austrian butterfly species, but a few difficult‐to‐identify species are not included or were grouped. Multiple models reached validation accuracies >97%, with larger models generally performing better, but there were some notable exceptions. The MaxViT‐tiny model achieved the highest accuracy despite being one of the smaller models. After further hyperparameter optimisation, the performance of the MaxViT was increased to a Top‐1 accuracy of 98.14% and a Top‐5 accuracy of 99.53% on test data. Although methods to correct for class imbalance were applied, recall and precision were generally lower for species with fewer images in the dataset. Our results emphasise the high potential of deep learning for species identification for the species analysed in this study. The results aid projects with limited computational capacities in choosing an efficient and appropriate model. Additional data processing steps, such as applying quality criteria to exclude model predictions with low confidence and expert assessments for difficult‐to‐determine species, can substantially increase the quality of identification.
No takes yet. Share an insight, caveat, or question.
Barkmann et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: