Image models usually VAE decode first, then super-res later, which is slow and mostly reconstructs rather than creates detail. This NVIDIA research makes the decoder itself a pixel diffusion model, so a 512^2
Fast and High-Resolution Latent Decoding with Pixel Diffusion
By
–
