Sodet-synthetic oversampling decision trees

Costa, Joaquim Fernando Pinto da; Alonso, Hugo

Sodet-synthetic oversampling decision trees

Ficheiros

Costa_et_al-2026-The_Journal_of_Supercomputing.pdf (367.23 KB)

Data

2026-04-19

Autores

Costa, Joaquim Fernando Pinto da

Alonso, Hugo

Editora

Springer

Idioma

Inglês

Resumo

In this work, we present a novel methodology for decision trees that use oversampling, not before tree construction (in the entire dataset), but inside each internal node (and corresponding input space region) of the tree. This strategy proves to be successful in fighting the greedy nature of decision trees. We take also into consideration the nature of the input variables, not just quantitative or binary, and also introduce the use of novel distances between instances that can also be used in other contexts. The application of our methodology to a significant number of datasets, thirteen, both balanced and imbalanced problems, shows the relevance of our approach when compared to CART and C5.0. Although our experiments were conducted on a standard computing platform, the proposed approach is well suited for high-performance computing environments, since node-level oversampling and distance computations can be efficiently parallelized, enabling the method to scale to large and high-dimensional datasets.

Palavras-chave

Decision trees, Oversampling, Synthetic samples, Data augmentation, Rare classes

Tipo de Documento

Artigo

Versão

VoR - Versão final publicada do editor

Versão da Editora

https://doi.org/10.1007/s11227-026-08503-8

Dataset

https://link.springer.com/article/10.1007/s11227-026-08503-8#citeas

Citação

Costa, J. F. P., & Alonso, H. (2026). Sodet-synthetic oversampling decision trees. Journal of Supercomputing, 82, 359, 1-33. https://doi.org/10.1007/s11227-026-08503-8. Repositório Institucional UPT. https://hdl.handle.net/11328/7088

Identificadores

https://hdl.handle.net/11328/7088
0920-8542
1573-0484

Tipo de Acesso

Acesso Aberto

Licença

Atribuição (CC-BY)

Ver registo completo