This work investigates adversarial robustness in deep learning models, focusing on the perceptual gap between cat classification and general object recognition. The study analyzes how adversarial perturbations affect model predictions and explores inconsistencies in learned representations. Experimental results highlight key robustness limitations and provide insights into improving model generalization and interpretability.
Indraneel Bose (Fri,) studied this question.