Abstract:Class imbalance is a frequently occurring scenario in classification tasks. Learning from imbalanced data poses a major challenge, which has instigated a lot of research in this area. Data preprocessing using sampling techniques is a standard approach to deal with the imbalance present in the data. Since standard classification algorithms do not perform well on imbalanced data, the dataset needs to be adequately balanced before training. This can be accomplished by oversampling the minority class or undersampling the majority class. In this study, a novel hybrid sampling algorithm has been proposed. To overcome the limitations of the sampling techniques while ensuring the quality of the retained sampled dataset, a sophisticated framework has been developed to properly combine three different sampling techniques. Neighborhood Cleaning rule is first applied to reduce the imbalance. Random undersampling is then strategically coupled with the SMOTE algorithm to obtain an optimal balance in the dataset. This proposed hybrid methodology, termed "SMOTE-RUS-NC", has been compared with other state-of-the-art sampling techniques. The strategy is further incorporated into the ensemble learning framework to obtain a more robust classification algorithm, termed "SRN-BRF". Rigorous experimentation has been conducted on 26 imbalanced datasets with varying degrees of imbalance. In virtually all datasets, the proposed two algorithms outperformed existing sampling strategies, in many cases by a substantial margin. Especially in highly imbalanced datasets where popular sampling techniques failed utterly, they achieved unparalleled performance. The superior results obtained demonstrate the efficacy of the proposed models and their potential to be powerful sampling algorithms in imbalanced domain.

Hybrid approaches for handling imbalanced structured and unstructured data

Addressing Binary Classification over Class Imbalanced Clinical Datasets Using Computationally Intelligent Techniques

Imbalanced Data Classification Algorithm Based on Integrated Sampling and Ensemble Learning.

Hybrid SVM algorithm oriented to classifying imbalanced datasets

A Survey of Methods for Managing the Classification and Solution of Data Imbalance Problem

Mitigating class imbalance in heart disease detection with machine learning

Handling Imbalanced Data: A Case Study for Binary Class Problems

Addressing Class Imbalance in Healthcare Data: Machine Learning Solutions for Age-Related Macular Degeneration and Preeclampsia

A Novel Imbalanced Data Classification Method Based on Weakly Supervised Learning for Fault Diagnosis

A literature survey on various aspect of class imbalance problem in data mining

Class Imbalance Problem in Data Mining Review

Optimizing Kernel Transformations to Handle Binary Class Imbalanced Dataset Classification

Effective treatment of imbalanced datasets in health care using modified SMOTE coupled with stacked deep learning algorithms

A hybrid ensemble and evolutionary algorithm for imbalanced classification and its application on bioinformatics

A hybridization of multiple imputation and one-class bagging ensemble approach for missing value and class imbalance problem

BALANCING SARCASTIC HINGLISH SHORT TEXT DATA USING AUGMENTATION TECHNIQUES WITH HANDLING SPELLING VARIATIONS

A Novel Hybrid Sampling Framework for Imbalanced Learning

Modeling of class imbalance handling with optimal deep learning enabled big data classification model

A Novel Adaptive Minority Oversampling Technique for Improved Classification in Data Imbalanced Scenarios

Hybrid Classifier for Optimizing Mental Health Prediction: Feature Engineering and Fusion Technique

Handling imbalanced medical datasets: review of a decade of research