What question did this study set out to answer?

This work aims to introduce a new decision tree methodology that incorporates oversampling within each internal node to improve classification accuracy.

April 21, 2026Open Access

Sodet—synthetic oversampling decision trees

Key Points

This work aims to introduce a new decision tree methodology that incorporates oversampling within each internal node to improve classification accuracy.
Developed a novel methodology for decision trees incorporating node-level oversampling.
Considered various types of input variables and introduced new distance metrics between instances.
Applied the methodology across thirteen datasets, both balanced and imbalanced.
The new approach showed significant improvements compared to CART and C5.0.
Demonstrated capability for efficient parallelization, making it scalable to large datasets.
Effectively addresses the limitations of traditional decision trees in greedy decision-making.

Abstract

Abstract In this work, we present a novel methodology for decision trees that use oversampling, not before tree construction (in the entire dataset), but inside each internal node (and corresponding input space region) of the tree. This strategy proves to be successful in fighting the greedy nature of decision trees. We take also into consideration the nature of the input variables, not just quantitative or binary, and also introduce the use of novel distances between instances that can also be used in other contexts. The application of our methodology to a significant number of datasets, thirteen, both balanced and imbalanced problems, shows the relevance of our approach when compared to CART and C5.0. Although our experiments were conducted on a standard computing platform, the proposed approach is well suited for high-performance computing environments, since node-level oversampling and distance computations can be efficiently parallelized, enabling the method to scale to large and high-dimensional datasets.

Bookmark

View Full Paper

Bookmark

View Full Paper

Sodet—synthetic oversampling decision trees

Key Points

Abstract

Cite This Study