Score-based generative models (SGMs) have demonstrated remarkable synthesis. SGMs rely on a diffusion process that gradually perturbs the data a tractable distribution, while the generative model learns to denoise. complexity of this denoising task is, apart from the data distribution, uniquely determined by the diffusion process. We argue that current employ overly simplistic diffusions, leading to unnecessarily complex processes, which limit generative modeling performance. Based on to statistical mechanics, we propose a novel critically-damped diffusion (CLD) and show that CLD-based SGMs achieve superior. CLD can be interpreted as running a joint diffusion in an extended, where the auxiliary variables can be considered "velocities" that are to the data variables as in Hamiltonian dynamics. We derive a novel matching objective for CLD and show that the model only needs to learn score function of the conditional distribution of the velocity given data, easier task than learning scores of the data directly. We also derive a new scheme for efficient synthesis from CLD-based diffusion models. We that CLD outperforms previous SGMs in synthesis quality for similar architectures and sampling compute budgets. We show that our novel for CLD significantly outperforms solvers such as Euler--Maruyama. Our provides new insights into score-based denoising diffusion models and be readily used for high-resolution image synthesis. Project page and code:://nv-tlabs.github.io/CLD-SGM.
No takes yet. Share an insight, caveat, or question.
Dockhorn et al. (2021) studied this question.