PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026IEEE Transactions on Pattern Analysis and Machine Intelligence2 citations

ZUMA: Training-free Zero-shot Unified Multimodal Anomaly Detection

View Full Paper
YMYunfeng MaMLMing LiuSJShuai Jiang

Key Points

  • The main goal is to develop a training-free framework for zero-shot multimodal anomaly detection that accommodates complex scenarios.
  • Introduced ZUMA, which utilizes CLIP for cross-modal anomaly detection without training.
  • Implemented cross-domain calibration to bridge the gap between 2D and 3D representations.
  • Employed dynamic semantic interaction to isolate anomaly regions using natural language as semantic anchors.
  • Developed ZUMA-FT, a fine-tuned version to assess performance improvements.
  • ZUMA set a new state-of-the-art performance in zero-shot multimodal anomaly detection on MVTec 3D-AD and Eyecandies.
  • ZUMA-FT achieved significant enhancements over ZUMA with minimal increase in parameters.
  • Demonstrated the capability of detecting anomalies across datasets and incomplete modalities.

Abstract

Multimodal anomaly detection (MAD) aims to exploit both texture and spatial attributes to identify deviations from normal patterns in complex scenarios. However, zero-shot (ZS) settings arising from privacy concerns or confidentiality constraints present significant challenges to existing MAD methods. To address this issue, we introduce ZUMA, a training-free, Zero-shot Unified Multimodal Anomaly detection framework that unleashes CLIP's cross-modal potential to perform ZS MAD. To mitigate the domain gap between CLIP's pretraining space and point clouds, we propose cross-domain calibration (CDC), which efficiently bridges the manifold misalignment through source-domain semantic transfer and establishes a hybrid semantic space, enabling a joint embedding of 2D and 3D representations. Subsequently, ZUMA performs dynamic semantic interaction (DSI) to enable structural decoupling of anomaly regions in the high-dimensional embedding space constructed by CDC, where natural languages serve as semantic anchors to help DSI establish discriminative hyperplanes within hybrid modality representations. Within this framework, ZUMA enables plug-and-play detection of 2D, 3D or multimodal anomalies, without training or fine-tuning even for cross-dataset or incomplete-modality scenarios. Additionally, to further investigate the potential of the training-free ZUMA within the training-based paradigm, we develop ZUMA-FT, a fine-tuned variant that achieves notable improvements with minimal parameter trade-off. Extensive experiments are conducted on two MAD benchmarks, MVTec 3D-AD and Eyecandies. Notably, the training-free ZUMA achieves state-of-the-art (SOTA) performance on both datasets, outperforming existing ZS MAD methods, including training-based approaches. Moreover, ZUMA-FT further extends the performance boundary of ZUMA with only 6.75 M learnable parameters. Code is available at: https://github.com/yif-ma/ZUMA.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ma et al. (2026) studied this question.

synapsesocial.com/papers/6980fdc7c1c9540dea80f701https://doi.org/10.1109/tpami.2026.3658856
Ask AI
Helpful
Bookmark
Share
View Full Paper