Abstract:BACKGROUND:Gene expression profiling has become a useful biological resource in recent years, and it plays an important role in a broad range of areas in biology. The raw gene expression data, usually in the form of large matrix, may contain missing values. The downstream analysis methods that postulate complete matrix input are thus not applicable. Several methods have been developed to solve this problem, such as K nearest neighbor impute method, Bayesian principal components analysis impute method, etc. In this paper, we introduce a novel imputing approach based on the Support Vector Regression (SVR) method. The proposed approach utilizes an orthogonal coding input scheme, which makes use of multi-missing values in one row of a certain gene expression profile and imputes the missing value into a much higher dimensional space, to obtain better performance.RESULTS:A comparative study of our method with the previously developed methods has been presented for the estimation of the missing values on six gene expression data sets. Among the three different input-vector coding schemes we tried, the orthogonal input coding scheme obtains the best estimation results with the minimum Normalized Root Mean Squared Error (NRMSE). The results also demonstrate that the SVR method has powerful estimation ability on different kinds of data sets with relatively small NRMSE.CONCLUSION:The SVR impute method shows better performance than, or at least comparable with, the previously developed methods in present research. The outstanding estimation ability of this impute method is partly due to the use of the most missing value information by incorporating orthogonal input coding scheme. In addition, the solid theoretical foundation of SVR method also helps in estimation of performance together with orthogonal input coding scheme. The promising estimation ability demonstrated in the results section suggests that the proposed approach provides a proper solution to the missing value estimation problem. The source code of the SVR method is available from http://202.38.78.189/downloads/svrimpute.html for non-commercial use.

Missing Value Estimation for DNA Microarray Gene Expression Data by Support Vector Regression Imputation and Orthogonal Coding Scheme

Novel Method for Missing Value Estimation in Gene Expression Profile Based on Support Vector Regression

Missing Microarray Data Estimation Based on Projection Onto Convex Sets Method

A hybrid imputation approach for microarray missing value estimation

Microarray Missing Data Imputation Based on A Set Theoretic Framework and Biological Constraints

A meta-data based method for DNA microarray imputation

DNA Microarray Data Imputation and Significance Analysis of Differential Expression

Missing value estimation for microarray data based on fuzzy C-means clustering

An efficient ensemble method for missing value imputation in microarray gene expression data

Missing value estimation for DNA microarray gene expression data with principal curves

Microarray Missing Value Imputation: A Regularized Local Learning Method

Missing Value Estimation Algorithms on Cluster and Representativeness Preservation of Gene Expression Microarray Data

Double Weighted Regression Estimation for Missing Values in Time Series Gene Expression Data

On-Line Imputation For Missing Values

Microarray Missing Value Imputation

Evaluations on Several Imputation Approaches of Integrated Omics Data

Usage of Clustering and Weighted Nearest Neighbors for Efficient Missing Data Imputation of Microarray Gene Expression Dataset

A Weighted Local Least Squares Imputation Method for Missing Value Estimation in Microarray Gene Expression Data.

Spanve: an Statistical Method to Detect Clustering-friendly Spatially Variable Genes in Large-scale Spatial Transcriptomics Data

A Global Learning with Local Preservation Method for Microarray Data Imputation

Effects of Replacing the Unreliable Cdna Microarray Measurements on the Disease Classification Based on Gene Expression Profiles and Functional Modules.