Experimental evaluation demonstrates improved visual inspection accuracy across industrial defect benchmarks, indicating the effectiveness of synthetic dialog fine-tuning for vision-language models.
Research on applying Vision-Language Models (VLM) to automate visual inspection has attracted growing attention. However, off-the-shelf VLMs lack the domain-specific expertise required to handle the wide variety of products and defects encountered in industrial inspection, making high-accuracy judgment difficult without additional training. In this study, we automatically construct dialog-style training data from a diverse set of images—including products not covered by existing evaluation benchmarks—using GPT-4o, and fine-tune a VLM (LLaVA-OneVision) with the generated data. Experimental results on the MMAD benchmark under zero-shot and one-shot settings show improved average accuracy across multiple tasks including good/bad decisions. In particular, we observe notable performance gains on MVTec AD and VisA.
No takes yet. Share an insight, caveat, or question.
Tsuchida et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: