Enhancing Imbalanced Data Classification with Class-Aware Synthetic Minority Over-sampling Technique
Ruturaj Mahajan, Sachin Patil, Vilabha PatilImbalanced data classification is one of the difficult machine learning tasks in which the majority class outnumbers the minority classes. This problem is prevalent in various domains such as medical disease detection, spam/fraud detection, digital marketing, agriculture and telecommunications. Synthetic Minority Over-sampling Technique (SMOTE) is a popular oversampling strategy that has been used in addressing the imbalanced dataset classification. However, SMOTE has limitations such as generating noisy samples, ignoring the underlying distribution of the data and over-sampling specific minority regions to the point of over-fitting. To address these limitations, several variants of SMOTE have been proposed, including ADASYN, Borderline-SMOTE, and Safe-Level-SMOTE. However, while these variants aim to address the limitations of SMOTE, they too have their own set of limitations, and their effectiveness may vary across different datasets and problem domains. Therefore, there is still room for improvement in over-sampling techniques to address the limitations of existing methods and improve their performance in diverse scenarios. As a solution, a novel extension of the SMOTE algorithm named "Class-Aware Synthetic Minority over-sampling Technique" (CA-SMOTE) has been proposed. By identifying and utilizing both minority and majority classes to create a synthetic sample. CA-SMOTE generates synthetic samples that are more diverse and closer to the underlying distribution of the data. With this approach, this technique aims to improve the accuracy and reliability of machine learning models trained on imbalanced data. The performance of CA-SMOTE is evaluated on several imbalanced datasets and compared it with SMOTE as well as other modern over-sampling algorithms. The experimental results prove that CA-SMOTE outperforms existing methods in terms of F1-score and AUC-ROC. Overall, CA-SMOTE can be a promising solution for enhanced imbalanced data classification in various real-world applications, improving the accuracy and reliability of machine learning models.