Leveraging epigenetic signatures to determine the cell-type of origin from long read sequencing data

Eilis Hannon,Jonathan Mill
DOI: https://doi.org/10.1101/2024.06.03.597114
2024-06-03
Abstract:DNA methylation differs across tissue- and cell-types with important implications for the analysis of disease-associated differences in tissues such as blood. To uncover the biological processes affected by epigenetic dysregulation, it is essential for epigenetic studies to generate data from the appropriate cell-types. Here we propose a framework to do this computationally from long-read sequencing data, bypassing the need to isolate subtypes of cells experimentally. Using reference data for six common blood cell-types, we evaluate the potential of this approach for attributing reads to specific cells using sequencing data generated from whole blood. Our analyses show that cell-type can be accurately classified using small regions of the genome comparable in size to those generated by long-read sequencing platforms, although the accuracy of classification varies across different regions of the genome and between cell-types. We found that for approximately one third of the genome it is possible to accurately discriminate reads originating from lymphocytes and myeloid cells with the prediction of more specialised subtypes of blood cell-types also encouraging. Our approach provides an alternative computational method for generating cell-specific DNA methylation profiles for epigenetic epidemiology, accelerating our ability to reveal critical insights of the role of the epigenome in health and disease.
Bioinformatics
What problem does this paper attempt to address?