Computational evaluation demonstrates enhanced adversarial robustness in large vision-language models via edge-guided prompting, indicating an efficient training-free defense mechanism.
Key Points
To establish an efficient, training-free defense mechanism that mitigates visual adversarial attacks against large vision-language models during inference.
Extracted structural edge maps from input images using the Canny edge detection operator.
Generated descriptions from edge representations using vision-language models and converted critical keywords into auxiliary text prompts.
Evaluated the defense framework across image classification and captioning tasks under three distinct visual adversarial attack settings.
Observed that structural image edge maps remain stable under adversarial perturbations while preserving core visual semantics.
Demonstrated that edge-guided keyword prompting significantly improves vision-language model robustness against multiple adversarial attack types without retraining parameters.