Deploying conditional Generative Adversarial Networks (cGANs) for industrial texture synthesis faces two barriers: the prohibitive cost of manual data annotation and the uncertain alignment between automated evaluation metrics and human perception. This study addresses both challenges for marble texture synthesis using 289 high-resolution industrial scans. We adapt an unsupervised segmentation pipeline combining Simple Linear Iterative Clustering (SLIC) superpixels, Gaussian Mixture Models (GMMs), and graph cut optimization to extract vein structures without manual annotation. Four cGAN architectures—baseline cGAN, Pix2Pix, BicycleGAN, and GauGAN—are benchmarked using a dual-evaluation protocol contrasting ten automated metrics with structured human-centered assessment. The results reveal a significant metric–perception discrepancy. Pix2Pix achieved the best Fréchet Inception Distance (FID = 85.3) yet received the lowest human ratings due to periodic texture artifacts. GauGAN produced textures statistically indistinguishable from real marble, achieving a Visual Turing Pass Rate (VTPR) of 0.533 and a Mean Opinion Score on Marble Authenticity (MOS-MA) of 2.89, despite an inferior FID (87.3). These findings make three contributions: an annotation-free segmentation pipeline, empirical evidence that automated metrics alone are insufficient for architecture selection, and a dual-evaluation framework that establishes human-in-the-loop assessment as essential for quality-critical industrial deployment.
Campos et al. (Tue,) studied this question.