PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 2026Computer Graphics Forum0 citationsOpen Access

MultiCOIN: Multi‐Modal COntrollable INbetweening

View Full Paper
MTM. TanveerYZY. ZhouSNS. Niklaus

Key Points

  • To develop a video inbetweening framework that allows enhanced control over frame transitions while ensuring quality video outputs.
  • Introduced MultiCOIN framework with multi-modal controls for video transitions.
  • Employed a Diffusion Transformer (DiT) for high-quality long video generation.
  • Separated content and motion into two branches for specialized processing.
  • Utilized a stage-wise training strategy for stable learning of controls.
  • Demonstrated improved motion complexity in video transitions.
  • Enhanced controllability over intermediate frames compared to existing methods.
  • Achieved greater narrative consistency throughout generated videos.

Abstract

Abstract Video inbetweening creates smooth transitions between two frames making it an indispensable tool for video editing and longform video synthesis. Existing methods struggle with large or complex motion and offer limited control over intermediate frames, often misaligning with user intent. We introduce MultiCOIN, a video inbetweening framework supporting multi‐modal controls, including depth transitions and layering, motion trajectories, text prompts, and target regions for movement localization. It balances flexibility, usability, and fine‐grained precision. Built on a Diffusion Transformer (DiT), due to its proven capability to generate high‐quality long video, our model maps all motion controls into a unified sparse point‐based representation compatible with the denoising process. Further, to respect the variety of controls which operate at varying levels of granularity and influence, we separate content and motion into two branches, enabling dedicated generators for each. A stage‐wise training strategy ensures stable learning of multi‐modal controls. Extensive experiments show improved motion complexity, controllability, and narrative consistency. Project Page: MultiCOIN.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tanveer et al. (2026) studied this question.

synapsesocial.com/papers/69e07d8f2f7e8953b7cbe7e1https://doi.org/10.1111/cgf.70362
Ask AI
Helpful
Bookmark
Share
View Full Paper