This survey reviews compression and quantization strategies for generative diffusion models, highlighting their implications for big data efficiency.
Generative diffusion models have emerged as a paradigmshifting approach in the field of artificial intelligence, demonstrating unprecedented capabilities in synthesizing high-fidelity, diverse, and multimodal data across a range of modalities including images, audio, and text. Despite their remarkable generative performance, these models are characterized by extremely large parameter counts, high computational complexity, and significant memory demands, which pose substantial challenges for deployment in real-world, large-scale, or resource-constrained environments. To address these limitations, a growing body of research has focused on model compression and quantization techniques, aiming to reduce memory footprint, accelerate inference, and enable efficient deployment on hardware with limited precision capabilities, without sacrificing generative fidelity. This work provides a comprehensive and detailed survey of compression and quantization strategies tailored specifically for generative diffusion models in the context of big data, where billions of samples may need to be processed or synthesized efficiently. We begin by presenting the mathematical foundations of diffusion processes, including the forward and reverse stochastic processes, iterative denoising mechanisms, and the sensitivity of these models to numerical perturbations, which underpin the critical challenges in applying compression and quantization. We then introduce a systematic taxonomy of methods, encompassing pruning, post-training quantization, quantization-aware training, mixed-precision schemes, low-rank factorization, knowledge distillation, and hybrid approaches, highlighting their principles, advantages, and limitations in the context of generative diffusion. Implementation strategies and hardware-aware considerations are discussed in depth, emphasizing memory hierarchy optimization, vectorized computation, structured sparsity, and adaptive precision allocation to maximize throughput and efficiency, particularly in distributed and large-scale data environments. Evaluation metrics are analyzed extensively, ranging from traditional generative quality measures such as Fréchet Inception Distance and Kernel Inception Distance to systemlevel metrics capturing latency, memory consumption, throughput, and robustness under iterative denoising. We further explore the key challenges and open research directions, including error propagation across iterative steps, adaptive and context-aware precision strategies, hybrid compression techniques, theoretical analyses of compression limits, and the development of large-scale benchmarking frameworks for big data. Finally, we outline emerging trends and future directions, emphasizing the integration of hardware-aware architecture search, dynamic quantization, co-design of compression strategies, and rigorous evaluation frameworks, which collectively aim to unlock the full potential of generative diffusion models in scalable, efficient, and robust AI systems. This survey provides a holistic reference for researchers and practitioners, synthesizing existing knowledge while identifying critical gaps and promising avenues for advancing the field of compressed and quantized generative diffusion models in the era of big data.
No takes yet. Share an insight, caveat, or question.
Jensen et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: