PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 2025ETRI Journal4 citationsOpen Access

High‐speed and precise virtual try‐on with two‐stage semantic segmentation and a latent consistency model for optimized diffusion processes

View Full Paper
SBSangyeop BaekJLJong Taek Lee

Key Points

  • HSP-VTON demonstrates a 2.8% improvement in mean intersection over union, enhancing segmentation precision.
  • Experimental results show that HSP-VTON outperforms existing methods on the ATR dataset and VITON-HD dataset.
  • The integration of a latent consistency model accelerates diffusion-based image generation without sacrificing quality.
  • By optimizing segmentation and generation, HSP-VTON addresses critical challenges in both precision and speed for virtual try-on.

Abstract

Abstract This work tests the hypothesis that the primary bottleneck for visual quality in virtual try‐on (VTON) systems is the precision of input segmentation masks, rather than generative capability. VTON technology empowers users to dress digital models in desired clothing items virtually. Conventional VTON models rely on segmentation models to isolate clothing regions and diffusion models to synthesize complete VTON images. This paper introduces high‐speed and precise VTON (HSP‐VTON) as a framework that uniquely combines refined two‐stage semantic segmentation for enhanced accuracy with a latent consistency model to accelerate the diffusion‐based image generation process. The synergistic integration of these components for VTON addresses critical challenges in both precision and speed. Experimental results on the ATR dataset demonstrate a 2.8% improvement in mean intersection over union compared with existing methods. Furthermore, HSP‐VTON achieves superior performance on the VITON‐HD dataset, outperforming state‐of‐the‐art VTON models. The latent consistency model also reduces the number of inference steps, leading to substantial time savings without compromising image quality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Baek et al. (2025) studied this question.

synapsesocial.com/papers/68e6a0f4718ef0a556b33dffhttps://doi.org/10.4218/etrij.2024-0592
Ask AI
Helpful
Bookmark
Share
View Full Paper