Performance of deep-learning based approaches to improve polygenic scores

Martin Kelemen,Yu Xu,Tao Jiang,Jing Hua Zhao,Carl Anderson,Chris Wallace,Adam S Butterworth,Michael Inouye
DOI: https://doi.org/10.1101/2024.10.23.24315973
2024-10-23
Abstract:Background/Objectives: Polygenic scores (PGS), which estimate an individual's genetic propensity for a disease or trait, have the potential to become part of genomic healthcare. In maximising the predictive performance of PGS, neural-network (NN) based deep learning has emerged as a method of intense interest to model complex, nonlinear phenomena, which may be adapted to exploit gene-gene (GxG) and gene-environment (GxE) interactions. Methods: To infer the amount of nonlinearity present in a phenotype, we present a framework for using NNs, which controls for the potential confounding effect of correlation between genetic variants, i.e. linkage disequilibrium (LD). We fit NN models to both simulated traits and 28 real disease and anthropometric traits in the UK Biobank. Results: Simulations confirmed that our framework adequately controls LD and can infer nonlinear effects, when such effects genuinely exist. Using this approach on real data, we found evidence for small amounts of nonlinearity due to GxG and GxE which mildly improved prediction performance (r2) by ~7% and ~4%, respectively. Despite evidence for nonlinear effects, NN models were outperformed by linear regression models for both genetic-only and genetic+environmental input scenarios with ~7% and ~5% differences in r2, respectively. Importantly, we found substantial evidence for confounding by joint tagging effects, whereby inferred GxG was actually LD with due to unaccounted for additive genetic variants. Conclusion: Our results indicate that the usefulness of NNs for generating polygenic scores for common traits and diseases may currently be limited and may be confounded by joint tagging effects due to LD.
Genetic and Genomic Medicine
What problem does this paper attempt to address?