This exploration combines CLIP and StyleGAN2 to improve text-driven facial image editing in sparse spaces, suggesting enhanced image diversity.
Due to the development of GAN and the proposal of many excellent models like StyleGAN, text-driven image editing and image generation have made great progress in recent years, but the task of generating diverse images of specific people under the guidance of text is still lacking. This paper combines two pre-training models, CLIP and StyleGAN2, to conduct a preliminary exploration of the above tasks. The latent code of the input portrait is driven to be edited and manipulated in the StyleGAN latent space via a CLIP-based text-driven module. Especially in the sparse region of the generator latent space, and when editing multiple attributes at the same time, some good results have finally been achieved.
No takes yet. Share an insight, caveat, or question.
Jianpeng Zou (2023) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: