We present a hierarchical VAE that, for the first time, generates samples while outperforming the PixelCNN in log-likelihood on all natural image. We begin by observing that, in theory, VAEs can actually represent models, as well as faster, better models if they exist, when sufficiently deep. Despite this, autoregressive models have historically VAEs in log-likelihood. We test if insufficient depth explains why scaling a VAE to greater stochastic depth than previously explored and it CIFAR-10, ImageNet, and FFHQ. In comparison to the PixelCNN, very deep VAEs achieve higher likelihoods, use fewer parameters, generate thousands of times faster, and are more easily applied to-resolution images. Qualitative studies suggest this is because the VAE efficient hierarchical visual representations. We release our source and models at https://github.com/openai/vdvae.
No takes yet. Share an insight, caveat, or question.
Rewon Child (2020) studied this question.