Recent success in deep learning has partially been driven by training overparametrized networks on ever larger datasets. It is therefore to ask: how much of the data is superfluous, which examples are for generalization, and how do we find them? In this work, we make striking observation that, in standard vision datasets, simple scores over several weight initializations can be used to identify important very early in training. We propose two such scores -- the Gradient (GraNd) and the Error L2-Norm (EL2N) scores -- and demonstrate their on a range of architectures and datasets by pruning significant of training data without sacrificing test accuracy. In fact, using2N scores calculated a few epochs into training, we can prune half of the10 training set while slightly improving test accuracy. Furthermore, for a dataset, EL2N scores from one architecture or hyperparameter generalize to other configurations. Compared to recent work that data by discarding examples that are rarely forgotten over the course of, our scores use only local information early in training. We also use scores to detect noisy examples and study training dynamics through the of important examples -- we investigate how the data distribution shapes loss surface and identify subspaces of the model's data representation that relatively stable over training.
No takes yet. Share an insight, caveat, or question.
Paul et al. (2021) studied this question.