PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 13, 20250 citationsOpen Access

Fine-Grained Controllable Apparel Showcase Image Generation via Garment-Centric Outpainting

View Full Paper
RZRong ZhangQinghai UniversityJWJingnan WangFirst Affiliated Hospital of Bengbu Medical CollegeZZZhiwen ZuoZhejiang Gongshang University

Key Points

  • The framework produces high-quality apparel showcase images while allowing for detailed customization based on text prompts.
  • A garment-adaptive pose prediction model is used to create diverse poses for the given garment, enhancing visual variety.
  • The multi-scale appearance customization module enables both overall and fine-grained control of the avatar's look during image generation.
  • Experiments show that this method outperforms existing techniques, ensuring better garment detail retention and efficiency.

Abstract

In this paper, we propose a novel garment-centric outpainting (GCO) framework based on the latent diffusion model (LDM) for fine-grained controllable apparel showcase image generation. The proposed framework aims at customizing a fashion model wearing a given garment via text prompts and facial images. Different from existing methods, our framework takes a garment image segmented from a dressed mannequin or a person as the input, eliminating the need for learning cloth deformation and ensuring faithful preservation of garment details. The proposed framework consists of two stages. In the first stage, we introduce a garment-adaptive pose prediction model that generates diverse poses given the garment. Then, in the next stage, we generate apparel showcase images, conditioned on the garment and the predicted poses, along with specified text prompts and facial images. Notably, a multi-scale appearance customization module (MS-ACM) is designed to allow both overall and fine-grained text-based control over the generated model's appearance. Moreover, we leverage a lightweight feature fusion operation without introducing any extra encoders or modules to integrate multiple conditions, which is more efficient. Extensive experiments validate the superior performance of our framework compared to state-of-the-art methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2025) studied this question.

synapsesocial.com/papers/68ece2abd1bb2827d1297207https://doi.org/10.48550/arxiv.2503.01294
Ask AI
Helpful
Bookmark
Share
View Full Paper