Identifying hierarchical cell states and gene signatures with deep exponential families for single-cell transcriptomics

Pedro F. Ferreira,Jack Kuipers,Niko Beerenwinkel
DOI: https://doi.org/10.1101/2022.10.15.512383
2024-01-04
Abstract:Single-cell gene expression data characterizes the complex heterogeneity of living systems. Tissues are composed of various cells with diverse cell states driven by different sets of genes. Cell states are often related in a hierarchical fashion, for example, in cell differentiation hierarchies. Clustering which respects a hierarchy, therefore, can improve functional interpretation and be leveraged to remove noise and batch effects when inferring gene signatures. For this task, we present single-cell Deep Exponential Families (scDEF), a multi-level Bayesian matrix factorization model for single-cell RNA-sequencing data. The model can identify hierarchies of cell states and be used for dimension reduction, gene signature identification, and batch integration. Additionally, it can be guided by known gene sets to jointly type cells and identify their hierarchical structure, or to find higher resolution states within the provided ones. In simulated and real data, scDEF outperforms alternative methods in finding cell populations across biologically distinct batches. We show that scDEF recovers cell type hierarchies in a whole adult animal, identifies a signature of response to interferon stimulation in peripheral blood mononuclear cells, and finds both patient-specific and shared cell states across nine high-grade serous ovarian cancer patients.
Bioinformatics
What problem does this paper attempt to address?