Randomized trial demonstrates enhanced few-shot classification in images, indicating improved data generation methods.
Generative data augmentation based on diffusion models has emerged as a promising approach for few-shot image classification. Existing methods, such as DA-Fusion, typically follow a “generate-once, use-directly” paradigm, which often suffers from uncontrollable generation quality, unstable semantic consistency, insufficient global diversity, and high sample redundancy. To address these limitations, we propose a two-stage Propose-and-Select framework for controllable data augmentation. This framework curates high-quality synthetic data offline, ensuring that no additional training overhead is introduced to downstream models. For selector optimization, our method eliminates the need for additional human annotations by leveraging the zero-shot prior knowledge of a vision–language model (CLIP) to construct relative-quality pseudo-labels. Furthermore, we develop an adaptive-temperature listwise ranking distillation objective to transfer quality-aware supervision effectively. We also introduce a multi-objective consistency regularization strategy to stabilize training and improve convergence. Under a strictly controlled augmentation budget, where all methods are provided with the same number of synthetic samples, the proposed approach consistently outperforms existing diffusion-based augmentation baselines across both few-shot classification benchmarks, achieving an accuracy of 79.58% on PASCAL VOC and 80.74% on the fine-grained Oxford 102 Flowers dataset. These results demonstrate the effectiveness of the proposed generation-selection paradigm in improving the quality, diversity, and semantic relevance of synthetic samples, thereby enhancing downstream few-shot classification performance.
No takes yet. Share an insight, caveat, or question.
Liu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: