Systematic review reveals efficiency-fidelity trade-offs in multimodal video synthesis models, indicating unresolved challenges in human motion plausibility and identity preservation.
Multimodal video diffusion models have emerged as transformative tools for controlled video synthesis, integrating text, images, audio, and pose sequences to generate semantically meaningful content. Despite significant advances, critical gaps persist in temporal consistency, multimodal alignment, and human-centric motion generation. Existing surveys have not addressed clearly the complex interplay between these components, particularly physiological constraints and identity preservation in human motion synthesis. This survey provides a comprehensive analysis through a unified architectural framework, examining spatial-temporal representations and multimodal conditioning mechanisms. We present the first systematic evaluation of human-centric motion modeling, addressing physiological plausibility and identity consistency challenges. Our analysis reveals fundamental trade-offs between computational efficiency and generation quality, with reported specialized techniques such as temporal block pruning achieving up to 523× computational savings under specific baselines with minimal quality degradation (see Sect. 6.3 for comparability caveats). Key findings indicate that current approaches struggle with seamless multimodal integration, human-centric applications face "uncanny valley" effects when physics constraints are too rigid, and identity preservation conflicts with motion dynamics. We introduce MIME-Vid (Multi-modal Integration with Motion Enhancement for Video Generation) as a conceptual reference framework that operationalises the unifying principles identified in our analysis (reference-flexibility, physics-perception asymmetry, and hierarchical disentanglement); empirical validation of MIME-Vid is deferred to follow-up work. Furthermore, we propose novel evaluation paradigms and identify future research directions for advancing multimodal video generation.
No takes yet. Share an insight, caveat, or question.
Albaghdadi et al. (2026) studied this question.