DOI: 10.35377/saucis...1833058 ISSN: 2636-8129

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Necati Vardar, Mehmet Fatih Ören
Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rarely systematically explored. This study presents a comprehensive comparative analysis of four classification algorithms—Naive Bayes (NB), K-Nearest Neighbors (K-NN), Artificial Neural Networks (ANN), and Random Forest (RF)—across ten benchmark datasets from the UCI Machine Learning Repository. Unlike previous studies relying on default parameters, this research employs a rigorous Grid Search strategy to optimize hyperparameters for each algorithm within a rigorous SMOTE-balanced stratified cross-validation pipeline to ensure robust evaluation. Performance was assessed using a wide range of metrics, including Accuracy, Precision, Recall, F1-score, and Area Under the Curve (AUC). Experimental results reveal that ANN achieved the highest robustness in high-dimensional and complex categorical datasets (e.g., Car Evaluation F1-score: 0.990), significantly outperforming traditional models. Conversely, RF demonstrated superior stability in datasets with high feature dimensionality (e.g., Arrhythmia F1-score: 0.600) and chemical interactions (e.g., QSAR Fish Toxicity F1-score: 0.839). While K-NN remained competitive in low-dimensional spaces, NB struggled with complex feature dependencies. This study contributes to the literature by demonstrating that algorithmic superiority is context-dependent and providing a data-driven framework for selecting classifiers based on structural characteristics such as dimensionality, categorical complexity, and sample size.