Abstract:Code smells indicate potential symptoms or problems in software due to inefficient design or incomplete implementation. These problems can affect software quality in the long-term. Code smell detection is fundamental to improving software quality and maintainability, reducing software failure risk, and helping to refactor the code. Previous works have applied several prediction methods for code smell detection. However, many of them show that machine learning (ML) and deep learning (DL) techniques are not always suitable for code smell detection due to the problem of imbalanced data. So, data imbalance is the main challenge for ML and DL techniques in detecting code smells. To overcome these challenges, this study aims to present a method for detecting code smell based on DL algorithms (Bidirectional Long Short-Term Memory (Bi-LSTM) and Gated Recurrent Unit (GRU)) combined with data balancing techniques (random oversampling and Tomek links) to mitigate data imbalance issue. To establish the effectiveness of the proposed models, the experiments were conducted on four code smells datasets (God class, data Class, feature envy, and long method) extracted from 74 open-source systems. We compare and evaluate the performance of the models according to seven different performance measures accuracy, precision, recall, f -measure, Matthew's correlation coefficient (MCC), the area under a receiver operating characteristic curve (AUC), the area under the precision–recall curve (AUCPR) and mean square error (MSE). After comparing the results obtained by the proposed models on the original and balanced data sets, we found out that the best accuracy of 98% was obtained for the Long method by using both models (Bi-LSTM and GRU) on the original datasets, the best accuracy of 100% was obtained for the long method by using both models (Bi-LSTM and GRU) on the balanced datasets (using random oversampling), and the best accuracy 99% was obtained for the long method by using Bi-LSTM model and 99% was obtained for the data class and Feature envy by using GRU model on the balanced datasets (using Tomek links). The results indicate that the use of data balancing techniques had a positive effect on the predictive accuracy of the models presented. The results show that the proposed models can detect the code smells more accurately and effectively.

CBReT: A Cluster-Based Resampling Technique for dealing with imbalanced data in code smell prediction

Improving accuracy of code smells detection using machine learning with data balancing techniques

A study of dealing class imbalance problem with machine learning methods for code smell severity detection using PCA-based feature selection technique

SMOTE-RkNN: A hybrid re-sampling method based on SMOTE and reverse k-nearest neighbors

A cluster-based SMOTE both-sampling (CSBBoost) ensemble algorithm for classifying imbalanced data

A Novel Adaptive Minority Oversampling Technique for Improved Classification in Data Imbalanced Scenarios

Do we need rebalancing strategies? A theoretical and empirical study around SMOTE and its variants

A Classfication Method For Imbalance Data Set Based on Kernel SMOTE

Over-sampling algorithm for imbalanced data classification

SMOTE: Synthetic Minority Over-sampling Technique

Alleviating Class Imbalance Issue in Software Fault Prediction Using DBSCAN-Based Induced Graph Under-Sampling Method

ReMix: Calibrated Resampling for Class Imbalance in Deep learning

A Novel Resampling Technique for Imbalanced Dataset Optimization

Enhancing and improving the performance of imbalanced class data using novel GBO and SSG: A comparative analysis

Survival Prediction from Imbalance colorectal cancer dataset using hybrid sampling methods and tree-based classifiers

A Clustering-Based Resampling Technique with Cluster Structure Analysis for Software Defect Detection in Imbalanced Datasets

Rare Event Prediction Using Similarity Majority Under-Sampling Technique

A Novel Hybrid Sampling Framework for Imbalanced Learning

The Impact of Class Rebalancing Techniques on the Performance and Interpretation of Defect Prediction Models

A cluster impurity-based hybrid resampling for imbalanced classification problems

Deep convolutional neural network model for bad code smells detection based on oversampling method