Abstract:Genomic malformations are believed to be the driving factors of many diseases. Therefore, understanding the intrinsic mechanisms underlying the genome and informing clinical practices have become two important missions of large-scale genomic research. Recently, high-throughput molecular data have provided abundant information about the whole genome, and have popularized computational tools in genomics. However, traditional machine learning methodologies often suffer from strong limitations when dealing with high-throughput genomic data, because the latter are usually very high dimensional, highly heterogeneous, and can show complicated nonlinear effects. In this thesis, we present five new algorithms or models to address these challenges, each of which is applied to a specific genomic problem.Project 1 focuses on model selection in cancer diagnosis. We develop an efficient algorithm (ADMM-ENSVM) for the Elastic Net Support Vector Machine, which achieves simultaneous variable selection and max-margin classification. On a colon cancer diagnosis dataset, ADMM-ENSVM shows advantages over other SVM algorithms in terms of diagnostic accuracy, feature selection ability, and computational efficiency.Project 2 focuses on model selection in gene correlation analysis. We develop an efficient algorithm (SBLVGG) using the similar methodology as of ADMM-ENSVM for the Latent Variable Gaussian Graphical Model (LVGG). LVGG models the marginal concentration matrix of observed variables as a combination of a sparse matrix and a low rank one. Evaluated on a microarray dataset containing 6,316 genes, SBLVGG is notably faster than the state-of-the-art LVGG solver, and shows that most of the correlation among genes can be effectively explained by only tens of latent factors.Project 3 focuses on ensemble learning in cancer survival analysis. We develop a gradient boosting model (GBMCI), which does not explicitly assume particular forms of hazard functions, but trains an ensemble of regression trees to approximately optimize the concordance index. We benchmark the performance of GBMCI against several popular survival models on a large-scale breast cancer prognosis dataset. GBMCI consistently outperforms other methods based on a number of feature representations, which are heterogeneous and contain missing values. Project 4 focuses on deep learning in gene expression inference (GEIDN). GEIDN is a large-scale neural network, which can infer ~21k target genes jointly from ~1k landmark genes and can naturally capture hierarchical nonlinear interactions among genes. We deploy deep learning techniques (drop out, momentum training, GPU computing, etc.) to train GEIDN. On a dataset of ~129k complete human transcriptomes, GEIDN outperforms both k-nearest neighbor regression and linear regression in predicting >99.96% of the target genes. Moreover, increased network scales help to improve GEIDN, while increased training data benefits GEIDN more than other methods.Project 5 focuses on deep learning in annotating coding and noncoding genetic variants (DANN). DANN is a neural network to differentiate evolutionarily derived alleles from simulated ones with 949 highly heterogeneous features. It can capture nonlinear relationships among features. We train DANN with deep learning techniques like for GEIDN. DANN achieves a 18.90% relative reduction in the error rate and a 14.52% relative increase in the area under the curve over CADD, a state-of-the-art algorithm to annotate genetic variants based on the linear SVM.

Machine Learning in Genome‐wide Association Studies

Machine Learning to Advance Human Genome-Wide Association Studies

Revisiting Genome-Wide Association Studies from Statistical Modelling to Machine Learning.

Reaching the End-Game for GWAS: Machine Learning Approaches for the Prioritization of Complex Disease Loci

A comprehensive investigation of statistical and machine learning approaches for predicting complex human diseases on genomic variants

Enhancing genotype-phenotype association with optimized machine learning and biological enrichment methods

Machine learning methods applied to genotyping data capture interactions between single nucleotide variants in late onset Alzheimer's disease

Current Status of Studies on Genome-Wide Gene-Gene Interactions

Identify Predictive SNP Groups in Genome Wide Association Study: A Sparse Learning Approach.

Valid inference for machine learning-assisted genome-wide association studies

High-throughput and Efficient Multilocus Genome-Wide Association Study on Longitudinal Outcomes.

Machine Learning for Large-Scale Genomics: Algorithms, Models and Applications

Machine learning methods applied to classify complex diseases using genomic data

Identifying Genetic Risk Factors via Sparse Group Lasso with Group Graph Structure

Current Status of SNPs Interaction in Genome-Wide Association Study

A Knowledge-Based Weighting Framework to Boost the Power of Genome-Wide Association Studies

An overview of SNP interactions in genome-wide association studies.

A Novel Approach to Encode Two-Way Epistatic Interactions Between Single Nucleotide Polymorphisms

Integrative functional linear model for genome-wide association studies with multiple traits

A Review of Regression and Classification Techniques for Analysis of Common and Rare Variants and Gene-Environmental Factors

Genetic studies through the lens of gene networks