Survey analyzes safety risks and ethical implications of integrating vision-language models in autonomous systems, indicating the need for improved reliability.
Embodied intelligence (EI), integrating vision-language models (VLMs) with action-oriented capabilities, presents transformative potential for autonomous systems. However, deploying VLMs in safety-critical applications like self-driving cars and collaborative robots introduces significant challenges. Key concerns include adversarial attacks and content manipulation, through which maliciously altered visual or textual inputs could lead to incorrect interpretations and hazardous actions in EI systems. Furthermore, integrating VLMs into embodied agents amplifies these risks, as real-world physical interactions introduce complex safety-critical scenarios where errors in perception or decision-making can have immediate and severe consequences. While VLM integration amplifies safety concerns, these models simultaneously offer a foundation for defensive strategies aimed at enhancing the reliability of EI systems. Additionally, VLM-based EI systems raise critical ethical concerns regarding accountability, embedded biases and societal displacement. This review systematically analyzes current applications, safety risks, protection methods and ethical implications, while proposing research directions to advance trustworthy EI systems.
No takes yet. Share an insight, caveat, or question.
Duan et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: