Comprehensive review analyzes interactions of adversarial machine learning, reinforcement learning, and explainable AI, highlighting security challenges.
Key Points
This review aims to understand the interactions between adversarial machine learning, reinforcement learning, and explainable AI under threat conditions.
Systematic analysis of 207 studies following PRISMA 2020 guidelines.
Taxonomy of adversarial attacks across training and inference phases constructed.
Evaluation of defense mechanisms and robustness across surveyed literature.
RL attack agents achieve evasion rates of 74–97% against ML-based detectors.
RL-based defenses report robustness gains of up to 3× compared to static baselines.
Attribution methods produce unreliable explanations under adversarial manipulations.