Abstract:For the DNA microarray datasets, tumor classification based on gene expression profiles has drawn great attention, and gene selection plays a significant role in improving the classification performance of microarray data. In this study, an effective hybrid gene selection method based on ReliefF and Ant colony optimization (ACO) algorithm for tumor classification is proposed. First, for the ReliefF algorithm, the average distance among k nearest or k non-nearest neighbor samples are introduced to estimate the difference among samples, based on which the distances between the samples in the same class or the different classes are defined, and then it can more effectively evaluate the weight values of genes for samples. To obtain the stable results in emergencies, a distance coefficient is developed to construct a new formula of updating weight coefficient of genes to further reduce the instability during calculations. When decreasing the distance between the same samples and increasing the distance between the different samples, the weight division is more obvious. Thus, the ReliefF algorithm can be improved to reduce the initial dimensionality of gene expression datasets and obtain a candidate gene subset. Second, a new pruning rule is designed to reduce dimensionality and obtain a new candidate subset with the smaller number of genes. The probability formula of the next point in the path selected by the ants is presented to highlight the closeness of the correlation relationship between the reaction variables. To increase the pheromone concentration of important genes, a new phenotype updating formula of the ACO algorithm is adopted to prevent the pheromone left by the ants that are overwhelmed with time, and then the weight coefficients of the genes are applied here to eliminate the interference of difference data as much as possible. It follows that the improved ACO algorithm has the ability of the strong positive feedback, which quickly converges to an optimal solution through the accumulation and the updating of pheromone. Finally, by combining the improved ReliefF algorithm and the improved ACO method, a hybrid filter-wrapper-based gene selection algorithm called as RFACO-GS is proposed. The experimental results under several public gene expression datasets demonstrate that the proposed method is very effective, which can significantly reduce the dimensionality of gene expression datasets, and select the most relevant genes with high classification accuracy.

Gene Selection Algorithm Based on Correlation Analysis

A Cancer Gene Selection Algorithm Based on the K-S Test and CFS

Parameters Selection in Gene Selection Using Gaussian Kernel Support Vector Machines by Genetic Algorithm

Gene selection and classification for cancer microarray data based on machine learning and similarity measures

A correlation guided genetic algorithm and its application to feature selection

Gene selection for cancer classification using a hybrid of univariate and multivariate feature selection methods

Feature Genes Selection Using Supervised Locally Linear Embedding and Correlation Coefficient for Microarray Classification

A Feature Selection Method Based on Feature Grouping and Genetic Algorithm

Gene Selection Using Gaussian Kernel Support Vector Machine Based Recursive Feature Elimination with Adaptive Kernel Width Strategy

Support Vector Machine-Recursive Feature Elimination for Localized Feature Selection

Gene selection and sample classification on microarray data based on adaptive genetic algorithm/k-nearest neighbor method

Gene selection using independent variable group analysis for tumor classification

Gene Selection for Cancer Classification using Support Vector Machines

Gene Features Selection for Three-Class Disease Classification via Multiple Orthogonal Partial Least Square Discriminant Analysis and S-Plot Using Microarray Data

Gene selection for cancer identification: a decision tree model empowered by particle swarm optimization algorithm

A Hybrid Gene Selection Method Based on ReliefF and Ant Colony Optimization Algorithm for Tumor Classification

A two-stage gene selection scheme utilizing MRMR filter and GA wrapper

A two-stage sparse logistic regression for optimal gene selection in high-dimensional microarray data classification

A Novel Approach for Single Gene Selection Using Clustering and Dimensionality Reduction

Feature (gene) Selection in Gene Expression-Based Tumor Classification

Novel Feature Selection Algorithms Based on Crowding Distance and Pearson Correlation Coefficient