Abstract This work tests the hypothesis that the primary bottleneck for visual quality in virtual try‐on (VTON) systems is the precision of input segmentation masks, rather than generative capability. VTON technology empowers users to dress digital models in desired clothing items virtually. Conventional VTON models rely on segmentation models to isolate clothing regions and diffusion models to synthesize complete VTON images. This paper introduces high‐speed and precise VTON (HSP‐VTON) as a framework that uniquely combines refined two‐stage semantic segmentation for enhanced accuracy with a latent consistency model to accelerate the diffusion‐based image generation process. The synergistic integration of these components for VTON addresses critical challenges in both precision and speed. Experimental results on the ATR dataset demonstrate a 2.8% improvement in mean intersection over union compared with existing methods. Furthermore, HSP‐VTON achieves superior performance on the VITON‐HD dataset, outperforming state‐of‐the‐art VTON models. The latent consistency model also reduces the number of inference steps, leading to substantial time savings without compromising image quality.
Baek et al. (2025) studied this question.