Experimental analysis demonstrates that data sampling improves classifier performance across imbalanced datasets, indicating that sampling efficacy depends on learner type and targeted metrics.
We present a comprehensive suite of experimentation on the subject of learning from imbalanced data. When classes are imbalanced, many learning algorithms can suffer from the perspective of reduced performance. Can data sampling be used to improve the performance of learners built from imbalanced data? Is the effectiveness of sampling related to the type of learner? Do the results change if the objective is to optimize different performance metrics? We address these and other issues in this work, showing that sampling in many cases will improve classifier performance.
No takes yet. Share an insight, caveat, or question.
Hulse et al. (2007) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: