PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 4, 2026Applied Sciences0 citationsOpen Access

Adversarial Attacks Against World Models: Hallucination-Driven Policy Failure

View Full Paper
JZJ ZhangHTHao TanRLRuonan Li

Key Points

  • The aim is to explore the adversarial risks associated with world models, focusing on verification of robustness and identification of vulnerabilities in modeling.
  • Conducted white-box attack experiments to verify model robustness.
  • Developed a latent space attack to mislead policies using only the perception module.
  • Introduced a temporal enhancement attack to manipulate future decisions through single-frame perturbations.
  • Quantitatively established the model's defensive boundaries against adversarial attacks.
  • Demonstrated that the latent space attack misled policies effectively without full model access.
  • Showed that the temporal enhancement attack amplified perturbed frame effects on future actions.

Abstract

World models have demonstrated powerful environment modeling capabilities in scenarios such as autonomous driving and robotics, but their adversarial security issues remain underexplored, in particular, adversarial risk analysis of world models. To bridge this gap, we systematically reveal the adversarial risks of world models through two dimensions: fundamental robustness verification and spatio-temporal vulnerability exploration. Specifically, in the fundamental robustness verification, we quantitatively certify the model’s defensive boundaries via white-box attack experiments; in the spatial dimension, by decoupling model component dependencies, we design a spatial-oriented gray-box attack, Latent Space Attack, that misleads policies by perturbing only the perception module, overcoming traditional requirements for full model access; in the temporal dimension, we propose a novel temporally correlated attack method, Temporal Enhancement Attack, incorporating cross-frame adversarial loss. By leveraging the model’s temporal memory and predictive capabilities, this approach amplifies single-frame perturbations to influence future decisions. This work conducts the first verification of the adversarial robustness of world models, thereby pointing out a new direction for adversarial research in reinforcement learning systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/6a211689d499ed480b16f784https://doi.org/10.3390/app16115484
Ask AI
Helpful
Bookmark
Share
View Full Paper