Scsagan: A Scrna-Seq Data Imputation Method Based on Semi-Supervised Learning and Probabilistic Latent Semantic Analysis

Zehao Xiong,Xiangtao Chen,Jiawei Luo,Cong Shen,Zhongyuan Xu
DOI: https://doi.org/10.1109/bibm55620.2022.9995463
2022-01-01
Abstract:single-cell RNA-sequencing (scRNA-seq) technology can reveal cellular heterogeneity with high throughput and resolution, facilitating the profiling of single-cell transcriptomes. However, due to some experimental factors, a large number of missing values are generated in scRNA-seq data, which are called dropout events, and this phenomenon affects the downstream analysis. Imputation is an effective denoising method, but existing imputation methods still face a huge challenge: lack of interpretability. In this study, we propose single-cell Self-Attention Generative Adversarial Networks(scSAGAN), a semi-supervised imputation method for scRNA-seq data. scSAGAN mainly uses Semi-Supervised Learning (SSL) and Probabilistic Latent Semantic Analysis (PLSA), which can not only learn the potential characteristics of different types of cells but explain their imputation behavior. In clustering experiments, scSAGAN exhibits better clustering performance than all baselines on 7 datasets. Next, we interpret the imputation behavior of scSAGAN on datasets such as Alzheimer’s disease and find causative genes associated with the corresponding datasets. scSAGAN is currently an open-source method, available at https://github.com/zehaoxiongl23/scSAGAN.
What problem does this paper attempt to address?