Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets

Peiran Jiang,Jose Lugo-Martinez
DOI: https://doi.org/10.1101/2023.08.25.554762
2023-08-28
bioRxiv
Abstract:Protein pockets are essential for many proteins to carry out their functions. Locating and measuring protein pockets as well as studying the anatomy of pockets helps us further understand protein function. Most research studies focus on learning either local or global information from protein structures. However, there is a lack of studies that leverage the power of integrating both local and global representations of these structures. In this work, we combine topological data analysis (TDA) and geometric deep learning (GDL) to analyze the putative protein pockets of enzymes. TDA captures blueprints of the global topological invariant of protein pockets, whereas GDL decomposes the fingerprints to building blocks of these pockets. This integration of local and global views provides a comprehensive and complementary understanding of the protein structural motifs (niches for short) within protein pockets. We also analyze the distribution of the building blocks making up the pocket and profile the predictive power of coupling local and global representations for the task of discriminating between enzymes and non-enzymes. We demonstrate that our representation learning framework for macromolecules is particularly useful when the structure is known, and the scenarios heavily rely on local and global information.
What problem does this paper attempt to address?