Sodet-synthetic oversampling decision trees
Data
2026-04-19
Embargo
Orientador
Coorientador
Título da revista
ISSN da revista
Título do volume
Editora
Springer
Idioma
Inglês
Título Alternativo
Resumo
In this work, we present a novel methodology for decision trees that use oversampling, not before tree construction (in the entire dataset), but inside each internal node (and corresponding input space region) of the tree. This strategy proves to be successful in fighting the greedy nature of decision trees. We take also into consideration the nature of the input variables, not just quantitative or binary, and also introduce the use of novel distances between instances that can also be used in other contexts. The application of our methodology to a significant number of datasets, thirteen, both balanced and imbalanced problems, shows the relevance of our approach when compared to CART and C5.0. Although our experiments were conducted on a standard computing platform, the proposed approach is well suited for high-performance computing environments, since node-level oversampling and distance computations can be efficiently parallelized, enabling the method to scale to large and high-dimensional datasets.
Palavras-chave
Decision trees, Oversampling, Synthetic samples, Data augmentation, Rare classes
Tipo de Documento
Artigo
Versão da Editora
Citação
Costa, J. F. P., & Alonso, H. (2026). Sodet-synthetic oversampling decision trees. Journal of Supercomputing, 82, 359, 1-33. https://doi.org/10.1007/s11227-026-08503-8. Repositório Institucional UPT. https://hdl.handle.net/11328/7088
Identificadores
TID
Designação
Tipo de Acesso
Acesso Aberto