PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 6, 2026IEEE Transactions on Image Processing0 citations

Active Dataset Distillation via Dual-Space Informative Matching

View Full Paper
DQDing QiTongji UniversityJLJian LiTencent (China)SDShiqing DouTongji University

Key Points

  • The study aims to enhance neural network training by developing an active learning algorithm for dataset distillation.
  • Proposed Active Dataset Distillation via Dual-Space Informative Matching (ACDD) algorithm.
  • Utilized an active learning approach to dynamically select informative real data subsets.
  • Implemented two interconnected loops: the dual-space active loop (DAL) and the distillation loop.
  • ACDD enhances training efficiency and generalization compared to state-of-the-art (SOTA) methods.
  • Demonstrated reduced real dataset requirements to 20%-40% of the original dataset size.
  • Achieved superior performance across multiple benchmarks including SVHN, CIFAR-10, CIFAR-100, TinyImageNet, and ImageNet subset.

Abstract

Dataset distillation improves neural network training efficiency by compressing large real datasets into compact synthetic datasets. Existing methods typically optimize matching objectives, such as aligning gradients, features, and trajectories between the synthetic and original datasets to ensure the distilled data retains essential properties for model training. However, many of these approaches rely on predefined distillation pools to streamline the process or treat all real data points equally, overlooking the dynamic nature of the synthetic dataset's training requirements during optimization. To address these limitations, we propose Active Dataset Distillation via Dual-Space Informative Matching (ACDD), an active learning-based algorithm that dynamically selects the most informative real data subset to align with the synthetic dataset's evolving needs. By adaptively refining the distillation pool, ACDD enhances training efficiency and generalization while ensuring the synthetic dataset effectively captures the original data's key characteristics. ACDD operates through two interconnected loops: the dual-space active loop (DAL) and the distillation loop. DAL plays a key role by dynamically selecting samples that balance diversity and uncertainty, adding them to the target distillation pool to meet the evolving informational needs of the current distillation loop. As a result, ACDD enables the synthetic dataset to achieve superior performance compared to SOTA methods across multiple benchmarks, including SVHN, CIFAR-10, CIFAR-100, TinyImageNet, and ImageNet subset. Moreover, ACDD reduces the required real dataset to just 20%-40% of the original, demonstrating its efficiency and effectiveness in data distillation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Qi et al. (2026) studied this question.

synapsesocial.com/papers/69aa6ee2531e4c4a9ff5915bhttps://doi.org/10.1109/tip.2026.3666748
Ask AI
Helpful
Bookmark
Share
View Full Paper