Classification models have emerged as the primary tools for numerous automatic computer vision tasks. However, they are susceptible to adversarial attacks, that can be harmful, but can also be employed to protect private information from classification-powered threat models designed to extract data from images. Black-box attacks, where an attacker has no knowledge about the model, are the most challenging ones, but also the most realistic ones. The difficulty is furthermore increased when one intends to create high-resolution adversarial images. We introduce NbuGAN, a novel black-box attack that creates high-resolution adversarial images deceiving image classification models in the targeted scenario. NbuGAN is experimentally validated: with 100 clean high-resolution images, NbuGAN creates 4275 high-resolution adversarial images that deceive 12 classification models trained on ImageNet for several clean-target combinations and expectations. Its average success rate is up to Formula: see text, each high-resolution adversarial image being obtained in less than a minute on average. NbuGAN is compared to nine state-of-the-art black-box and white-box attacks. NbuGAN not only significantly outperforms the black-box attacks, but its remarkable speed, its success rates and the exceptional visual quality of the created high-resolution adversarial images, make NbuGAN highly competitive, both intrinsically and comparatively, even against white-box attacks.
Topal et al. (Mon,) studied this question.