3D genome reconstruction from partially phased Hi-C data

Diego Cifuentes,Jan Draisma,Oskar Henriksson,Annachiara Korchmaros,Kaie Kubjas
DOI: https://doi.org/10.1007/s11538-024-01263-7
2024-05-29
Abstract:The 3-dimensional (3D) structure of the genome is of significant importance for many cellular processes. In this paper, we study the problem of reconstructing the 3D structure of chromosomes from Hi-C data of diploid organisms, which poses additional challenges compared to the better-studied haploid setting. With the help of techniques from algebraic geometry, we prove that a small amount of phased data is sufficient to ensure finite identifiability, both for noiseless and noisy data. In the light of these results, we propose a new 3D reconstruction method based on semidefinite programming, paired with numerical algebraic geometry and local optimization. The performance of this method is tested on several simulated datasets under different noise levels and with different amounts of phased data. We also apply it to a real dataset from mouse X chromosomes, and we are then able to recover previously known structural features.
Genomics,Algebraic Geometry,Optimization and Control
What problem does this paper attempt to address?