Abstract:The genomic evaluation process relies on the assumption of linkage disequilibrium between dense single-nucleotide polymorphism (SNP) markers at the genome level and quantitative trait loci (QTL). The present study was conducted with the aim of evaluating four frequentist methods including Ridge Regression, Least Absolute Shrinkage and Selection Operator (LASSO), Elastic Net, and Genomic Best Linear Unbiased Prediction (GBLUP) and five Bayesian methods including Bayes Ridge Regression (BRR), Bayes A, Bayesian LASSO, Bayes C, and Bayes B, in genomic selection using simulation data. The difference between prediction accuracy was assessed in pairs based on statistical significance (p-value) (i.e., t test and Mann-Whitney U test) and practical significance (Cohen's d effect size) For this purpose, the data were simulated based on two scenarios in different marker densities (4000 and 8000, in the whole genome). The simulated data included a genome with four chromosomes, 1 Morgan each, on which 100 randomly distributed QTL and two different densities of evenly distributed SNPs (1000 and 2000), at the heritability level of 0.4, was considered. For the frequentist methods except for GBLUP, the regularization parameter λ was calculated using a five-fold cross-validation approach. For both scenarios, among the frequentist methods, the highest prediction accuracy was observed by Ridge Regression and GBLUP. The lowest and the highest bias were shown by Ridge Regression and GBLUP, respectively. Also, among the Bayesian methods, Bayes B and BRR showed the highest and lowest prediction accuracy, respectively. The lowest bias in both scenarios was registered by Bayesian LASSO and the highest bias in the first and the second scenario were shown by BRR and Bayes B, respectively. Across all the studied methods in both scenarios, the highest and the lowest accuracy were shown by Bayes B and LASSO and Elastic Net, respectively. As expected, the greatest similarity in performance was observed between GBLUP and BRR ( d = 0.007 , in the first scenario and d = 0.003 , in the second scenario). The results obtained from parametric t and non-parametric Mann-Whitney U tests were similar. In the first and second scenario, out of 36 t test between the performance of the studied methods in each scenario, 14 ( P < . 001 ) and 2 ( P < . 05 ) comparisons were significant, respectively, which indicates that with the increase in the number of predictors, the difference in the performance of different methods decreases. This was proven based on the Cohen's d effect size, so that with the increase in the complexity of the model, the effect size was not seen as very large. The regularization parameters in frequentist methods should be optimized by cross-validation approach before using these methods in genomic evaluation.

Comparing statistical learning methods for complex trait prediction from gene expression

Leveraging Multiple Layers of Data To Predict Drosophila Complex Traits

Nonlinear network-based quantitative trait prediction from transcriptomic data

Quantifying the Relationship Between Gene Expressions and Trait Values in General Pedigrees.

Integrating Mrna Transcripts and Genomic Information into Genomic Prediction.

Of mice and men: Sparse statistical modeling in cardiovascular genomics

A comprehensive investigation of statistical and machine learning approaches for predicting complex human diseases on genomic variants

On the cross-population generalizability of gene expression prediction models

Predicting the genetic component of gene expression using gene regulatory networks

A Comparative Analysis of Gene Expression Profiling by Statistical and Machine Learning Approaches

Incorporating Gene Annotation into Genomic Prediction of Complex Phenotypes

Standard machine learning approaches outperform deep representation learning on phenotype prediction from transcriptomics data

Integrative approaches for large-scale transcriptome-wide association studies

Investigating the Performance of Frequentist and Bayesian Techniques in Genomic Evaluation

Comparison of machine learning methods for genomic prediction of selected Arabidopsis thaliana traits

Statistical Learning Methods For Genome-based Analysis Of Quantitative Traits

Genetic risk prediction in complex disease

Predicting fitness related traits using gene expression and machine learning

Appraisal of Gene Expression-Based Classifiers for Neuropsychiatric Disorders: A Meta-Regression

Integrating Gene Expression Data into Single-Step Method (ssblup) Improves Genomic Prediction Accuracy for Complex Traits of Duroc × Erhualian F2 Pig Population

Trait genetic architecture and population structure determine model selection for genomic prediction in natural Arabidopsis Thaliana populations