A novel method for data fusion over entity-relation graphs and its application to protein-protein interaction prediction

Y. Moreau,D. Raimondi,Adam Arany,J. Simm
DOI: https://doi.org/10.1093/bioinformatics/btab092
2021-02-09
Abstract:MOTIVATION Modern Bioinformatics is facing increasingly complex problems to solve, and we are indeed rapidly approaching an era in which the ability to seamlessly integrate heterogeneous sources of information will be crucial for the scientific progress. Here we present a novel non-linear data fusion framework that generalizes the conventional Matrix Factorization paradigm allowing inference over arbitrary Entity-Relation graphs, and we applied it to the prediction of Protein-Protein Interactions (PPIs). Improving our knowledge of Protein Protein Interaction (PPI) networks at the proteome scale is indeed crucial to understand protein function, physiological and disease states and cell life in general. RESULTS We devised three data-fusion based models for the proteome-level prediction of PPIs, and we show that our method outperforms state of the art approaches on common benchmarks. Moreover, we investigate its predictions on newly published PPIs, showing that this new data has a clear shift in its underlying distributions and we thus train and test our models on this extended dataset. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.
Medicine,Computer Science,Biology
What problem does this paper attempt to address?