Contamde: Differential Expression Analysis of RNA-seq Data for Contaminated Tumor Samples.

Qi Shen,Jiyuan Hu,Ning Jiang,Xiaohua Hu,Zewei Luo,Hong Zhang
DOI: https://doi.org/10.1093/bioinformatics/btv657
IF: 5.8
2015-01-01
Bioinformatics
Abstract:MOTIVATIONAccurate detection of differentially expressed genes between tumor and normal samples is a primary approach of cancer-related biomarker identification. Due to the infiltration of tumor surrounding normal cells, the expression data derived from tumor samples would always be contaminated with normal cells. Ignoring such cellular contamination would deflate the power of detecting DE genes and further confound the biological interpretation of the analysis results. For the time being, there does not exists any differential expression analysis approach for RNA-seq data in literature that can properly account for the contamination of tumor samples.RESULTSWithout appealing to any extra information, we develop a new method 'contamDE' based on a novel statistical model that associates RNA-seq expression levels with cell types. It is demonstrated through simulation studies that contamDE could be much more powerful than the existing methods that ignore the contamination. In the application to two cancer studies, contamDE uniquely found several potential therapy and prognostic biomarkers of prostate cancer and non-small cell lung cancer.AVAILABILITY AND IMPLEMENTATIONAn R package contamDE is freely available at http://homepage.fudan.edu.cn/zhangh/softwares/CONTACTzhanghfd@fudan.edu.cnSUPPLEMENTARY INFORMATIONSupplementary data are available at Bioinformatics online.
What problem does this paper attempt to address?