Generative diffusion models have achieved substantial progress in image content creation, yet practical interactive generation remains challenging when heterogeneous conditions must be jointly interpreted, fine-grained attributes must be controlled independently, and users need to iteratively refine generated results. This paper proposes a Multimodal Interactive Generation Framework (MIGF) that integrates three task-oriented components: Adaptive Condition Fusion (ACF), Semantic Decoupling (SDM), and an Interactive Feedback Mechanism (IFM). ACF addresses sample-dependent reliability differences among text, reference-image, and structured conditions by predicting modality-specific fusion weights for each input. SDM uses contrastive supervision to organize latent representations into content- and style-related subspaces, facilitating more independent attribute control. IFM integrates local modification, attribute adjustment, and example guidance into a unified progressive refinement process. We construct a multimodal dataset containing 50,000 high-resolution images covering natural scenes, artistic works, design patterns, and architecture. Under matched evaluation settings, MIGF reduces FID by 23.7%, improves CLIP Score by 18.6%, and reduces LPIPS by 12.7% relative to the strongest self-run baseline, ControlNet. Ablation experiments show consistent contributions from the three components. The combination of text, image, and sketch conditions achieves a Control Precision of 0.876, corresponding to an 11.9% improvement over the strongest single-modality configuration. In addition, the fast inference mode reduces generation time to 1.9 s per image. These results indicate that the proposed framework provides a practical approach to multimodal and interactively controllable image generation, while its remaining limitations in robustness, computational cost, and cross-domain generalization motivate further investigation.
No takes yet. Share an insight, caveat, or question.
Wu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: