Abstract:Genomic data integration-the process of statistically combining diverse Sources of information from functional genomics experiments to make large-scale predictions-is becoming increasingly prevalent. One might expect that this process should become progressively more powerful With the integration of more evidence. Here, we explore the limits of genomic data integration, assessing the degree to which predictive power increases with the addition of more features. We focus oil a predictive context that has been extensively investigated and benchmarked in the past-the prediction of protein-protein interactions in yeast. We start by using a simple Naive Bayes classifier for integrating diverse Sources of genomic evidence, ranging from coexpression relationships to similar phylogenetic profiles. We expand the number of features considered for prediction to 16, significantly more than previous Studies. Overall, we observe a small, but measurable improvement in prediction performance over previous benchmarks, based on four strong features. This allows us to identify new yeast interactions with high confidence. It also allows us to quantitatively assess the inter-relations amongst different genomic features. It is known that subtle correlations and dependencies between features call confound the strength of interaction predictions. We investigate this issue in detail through calculating mutual information. To Our Surprise, we find no appreciable statistical dependence between the many possible pairs of features. We further explore feature dependencies by comparing the performance Of Our simple Naive Bayes classifier with a boosted version of the same classifier, which is fairly resistant to feature dependence. We find that boosting does not improve performance, indicating that, at least for prediction purposes, Our genomic features are essentially independent. In Summary, by integrating a few (i.e., four) good features, we approach the maximal predictive power of current genomic data integration; moreover, this limitation does not reflect (potentially removable) inter-relationships between the features.

Polygenic prediction and gene regulation networks

Towards Prediction and Prioritization of Disease Genes by the Modularity of Human Phenome-Genome Assembled Network.

Performance of deep-learning based approaches to improve polygenic scores

Influence of Genetic Interactions on Polygenic Prediction

Predicting missing expression values in gene regulatory networks using a discrete logic modeling optimization guided by network stable states

Deep learning for polygenic prediction: The role of heritability, interaction type and sample size

Learning a nonlinear dynamical system model of gene regulation: A perturbed steady-state approach

Modeling gene interactions in polygenic prediction via geometric deep learning

Predicting the genetic component of gene expression using gene regulatory networks

Nonlinear network-based quantitative trait prediction from transcriptomic data

Higher-order epistasis and phenotypic prediction

Dissecting Genetic Networks Underlying Complex Phenotypes: The Theoretical Framework

Inferring multilayer interactome networks shaping phenotypic plasticity and evolution

Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes

GENet: A Graph-Based Model Leveraging Histone Marks and Transcription Factors for Enhanced Gene Expression Prediction

Integrative Analysis of Genetical Genomics Data Incorporating Network Structures.

Advancing regulatory genomics with machine learning

A semisupervised model to predict regulatory effects of genetic variants at single nucleotide resolution using massively parallel reporter assays

Assessing The Limits Of Genomic Data Integration For Predicting Protein Networks

Genomic Prediction of Complex Phenotypes Using Genic Similarity Based Relatedness Matrix

Proteome-scale prediction of molecular mechanisms underlying dominant genetic diseases