PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 26, 20240 citationsOpen Access

User-Friendly Customized Generation with Multi-Modal Prompts

View Full Paper
ZLZhong Lin-HaoHYHong YanWCWentao Chen

Key Points

Key points are not available for this paper at this time.

Abstract

Text-to-image generation models have seen considerable advancement, catering to the increasing interest in personalized image creation. Current customization techniques often necessitate users to provide multiple images (typically 3-5) for each customized object, along with the classification of these objects and descriptive textual prompts for scenes. This paper questions whether the process can be made more user-friendly and the customization more intricate. We propose a method where users need only provide images along with text for each customization topic, and necessitates only a single image per visual concept. We introduce the concept of a ``multi-modal prompt'', a novel integration of text and images tailored to each customization concept, which simplifies user interaction and facilitates precise customization of both objects and scenes. Our proposed paradigm for customized text-to-image generation surpasses existing finetune-based methods in user-friendliness and the ability to customize complex objects with user-friendly inputs. Our code is available at https: //github. com/zhongzero/Multi-Modal-Prompthttps: //github. com/zhongzero/Multi-Modal-Prompt.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lin-Hao et al. (2024) studied this question.

synapsesocial.com/papers/68e685a5b6db64358760ee2dhttps://doi.org/10.48550/arxiv.2405.16501
Ask AI
Helpful
Bookmark
Share
View Full Paper