We propose EcoLoop, a framework for constructing domain-specific perceptual quality metrics for synthetic image generation via fine-tuned dense captioning models. The approach fine-tunes a vision-language model to produce exhaustive, expert-grade descriptions of domain imagery at multiple spatial scales, then uses these caption–image pairs to fine-tune a CLIP-family model into a domain-adapted embedding space. The resulting model serves as both a scoring function for synthetic image quality and a steering signal for conditional generation pipelines. We validate the approach in marine wildlife imagery (cetacean photography), where morphological features critical for species identification and health assessment are underrepresented in generic perceptual metrics. The framework generalizes to any domain where expert visual perception diverges from lay perception, including medical imaging, materials science, and remote sensing.
Dalal Aryaman (Wed,) studied this question.