This analysis demonstrates improved facial image generation using diffusion models and measurable metrics indicate better performance.
To solve the challenges such as high requirements for the technical skills of portrait artists and low drawing efficiency in the traditional simulated portrait technique faces. A multi-condition input strategy to improve the generation effect of the model was proposed based on the Brownian Bridge Diffusion Model (BBDM) framework. Specifically, the original Brownian bridge diffusion process was split into two parts by introducing new conditional information. The model was trained to extract more detailed information from sketches at the first part. Then, the model was trained to reconstruct more realistic faces by obtaining detailed information. Compared with the BBDM, the Fréchet Inception Distance (FID) of the proposed model increased from 18.745 to 14.440, and the Learned Perceptual Image Patch Similarity (LPIPS) decreased from 0.2937 to 0.2663. Competitive scores were obtained in the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM). Compared to the other methods based on the Generative Adversarial Networks (GANs), the FID, LPIPS, and SSIM scores were the highest, at 14.440, 0.2663, and 0.5530, respectively. The experimental results showed that our method generates more realistic facial details and achieved better performance in measurable metrics.
No takes yet. Share an insight, caveat, or question.
Li et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: