Abstract:Genomic malformations are believed to be the driving factors of many diseases. Therefore, understanding the intrinsic mechanisms underlying the genome and informing clinical practices have become two important missions of large-scale genomic research. Recently, high-throughput molecular data have provided abundant information about the whole genome, and have popularized computational tools in genomics. However, traditional machine learning methodologies often suffer from strong limitations when dealing with high-throughput genomic data, because the latter are usually very high dimensional, highly heterogeneous, and can show complicated nonlinear effects. In this thesis, we present five new algorithms or models to address these challenges, each of which is applied to a specific genomic problem.Project 1 focuses on model selection in cancer diagnosis. We develop an efficient algorithm (ADMM-ENSVM) for the Elastic Net Support Vector Machine, which achieves simultaneous variable selection and max-margin classification. On a colon cancer diagnosis dataset, ADMM-ENSVM shows advantages over other SVM algorithms in terms of diagnostic accuracy, feature selection ability, and computational efficiency.Project 2 focuses on model selection in gene correlation analysis. We develop an efficient algorithm (SBLVGG) using the similar methodology as of ADMM-ENSVM for the Latent Variable Gaussian Graphical Model (LVGG). LVGG models the marginal concentration matrix of observed variables as a combination of a sparse matrix and a low rank one. Evaluated on a microarray dataset containing 6,316 genes, SBLVGG is notably faster than the state-of-the-art LVGG solver, and shows that most of the correlation among genes can be effectively explained by only tens of latent factors.Project 3 focuses on ensemble learning in cancer survival analysis. We develop a gradient boosting model (GBMCI), which does not explicitly assume particular forms of hazard functions, but trains an ensemble of regression trees to approximately optimize the concordance index. We benchmark the performance of GBMCI against several popular survival models on a large-scale breast cancer prognosis dataset. GBMCI consistently outperforms other methods based on a number of feature representations, which are heterogeneous and contain missing values. Project 4 focuses on deep learning in gene expression inference (GEIDN). GEIDN is a large-scale neural network, which can infer ~21k target genes jointly from ~1k landmark genes and can naturally capture hierarchical nonlinear interactions among genes. We deploy deep learning techniques (drop out, momentum training, GPU computing, etc.) to train GEIDN. On a dataset of ~129k complete human transcriptomes, GEIDN outperforms both k-nearest neighbor regression and linear regression in predicting >99.96% of the target genes. Moreover, increased network scales help to improve GEIDN, while increased training data benefits GEIDN more than other methods.Project 5 focuses on deep learning in annotating coding and noncoding genetic variants (DANN). DANN is a neural network to differentiate evolutionarily derived alleles from simulated ones with 949 highly heterogeneous features. It can capture nonlinear relationships among features. We train DANN with deep learning techniques like for GEIDN. DANN achieves a 18.90% relative reduction in the error rate and a 14.52% relative increase in the area under the curve over CADD, a state-of-the-art algorithm to annotate genetic variants based on the linear SVM.

Navigating the pitfalls of applying machine learning in genomics

[Effects of some counter ions on physicochemical properties and the local anesthetic action of dibucaine base. 3. Effects on blocking of conduction in frog's sciatic nerve fibers and solvent affinity].

Interpretable Machine Learning for Genomics

What makes a good prediction? Feature importance and beginning to open the black box of machine learning in genetics

Machine Learning to Advance Human Genome-Wide Association Studies

Machine learning applications in genetics and genomics

Advancing regulatory genomics with machine learning

Machine Learning and Deep Learning in Genetics and Genomics

Machine Learning and Integrative Analysis of Biomedical Big Data

Machine Learning for Large-Scale Genomics: Algorithms, Models and Applications

Integration of Artificial Intelligence, Machine Learning and Deep Learning Techniques in Genomics: Review on Computational Perspectives for NGS Analysis of DNA and RNA Seq Data

Machine learning in onco-pharmacogenomics: a path to precision medicine with many challenges

Machine learning for integrating data in biology and medicine: Principles, practice, and opportunities

Biosystems Design by Machine Learning

N-Substituted pyrazole-3-carboxamides as inhibitors of human 15-lipoxygenase.

The era of big data: Genome-scale modelling meets machine learning

Machine learning in rare disease

Opening the Black Box: Interpretable Machine Learning for Geneticists.

Understanding Biology in the Age of Artificial Intelligence

Advances in AI and machine learning for predictive medicine

Machine learning and genomics: precision medicine vs. patient privacy