DOI: 10.35378/gujs.1643918 ISSN: 2147-1762

A Comparative Study of Data Balancing Algorithms to Examine Optimal Class Distribution in Imbalanced Datasets: A Simulation Based Approach

Hakan Öztürk, Mevlüt Türe, İmran Kurt Omurlu
In a two-class dataset, the class imbalance problem arises if there is a considerable difference between the number of samples in the classes. Many data balancing algorithms have been proposed to address this issue. However, only a limited number of studies have examined the candidate balance ratios of certain balancing algorithms, often focusing on real datasets. Unlike previous studies, this research evaluates seven balancing algorithms in terms of their predictive performance for an estimated population parameter (EP) and examines minority-majority class distributions yielding performance comparable to EP using an original simulation scenario. In this study, imbalanced datasets were sampled from a simulated population dataset and gradually balanced using random oversampling (ROS), synthetic minority oversampling technique (SMOTE), majority weighted minority oversampling technique (MWMOTE), adaptive synthetic sampling approach (ADASYN), random undersampling (RUS), random under boosting (RUSBoost), and under bagging (UB) algorithms. The classification and regression trees (CART) method was used to classify the data at each step, and the area under the ROC curve (AUC) was employed to evaluate the performance of the balancing algorithms. The findings obtained under the present simulation setting indicate that RUSBoost and UB algorithms yield statistically higher results than EP when certain balance ratios are exceeded. Meanwhile, within the evaluated simulation setting and the CART–AUC framework, other methods do not surpass EP and generally achieve their highest mean AUC values at full balance (50:50).

More from our Archive