Abstract:A novel framework for the automated evaluation of various deep learning-based splice site detectors is presented. The framework eliminates time-consuming development and experimenting activities for different codebases, architectures, and configurations to obtain the best models for a given RNA splice site dataset. RNA splicing is a cellular process in which pre-mRNAs are processed into mature mRNAs and used to produce multiple mRNA transcripts from a single gene sequence. Since the advancement of sequencing technologies, many splice site variants have been identified and associated with the diseases. So, RNA splice site prediction is essential for gene finding, genome annotation, disease-causing variants, and identification of potential biomarkers. Recently, deep learning models performed highly accurately for classifying genomic signals. Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) and its bidirectional version (BLSTM), Gated Recurrent Unit (GRU), and its bidirectional version (BGRU) are promising models. During genomic data analysis, CNN's locality feature helps where each nucleotide correlates with other bases in its vicinity. In contrast, BLSTM can be trained bidirectionally, allowing sequential data to be processed from forward and reverse directions. Therefore, it can process 1-D encoded genomic data effectively. Even though both methods have been used in the literature, a performance comparison was missing. To compare selected models under similar conditions, we have created a blueprint for a series of networks with five different levels. As a case study, we compared CNN and BLSTM models' learning capabilities as building blocks for RNA splice site prediction in two different datasets. Overall, CNN performed better with accuracy ( improvement), F1 score ( improvement), and AUC-PR ( improvement) in human splice site prediction. Likewise, an outperforming performance with accuracy ( improvement), F1 score ( improvement), and AUC-PR ( improvement) is achieved in C. elegans splice site prediction. Overall, our results showed that CNN learns faster than BLSTM and BGRU. Moreover, CNN performs better at extracting sequence patterns than BLSTM and BGRU. To our knowledge, no other framework is developed explicitly for evaluating splice detection models to decide the best possible model in an automated manner. So, the proposed framework and the blueprint would help selecting different deep learning models, such as CNN vs. BLSTM and BGRU, for splice site analysis or similar classification tasks and in different problems.

DeepSS: Exploring Splice Site Motif Through Convolutional Neural Network Directly from DNA Sequence

Identifying DNA Sequence Motifs Using Deep Learning

An effective deep learning-based approach for splice site identification in gene expression

DRANetSplicer: A Splice Site Prediction Model Based on Deep Residual Attention Networks

An automated framework for evaluation of deep learning models for splice site predictions

SpliceRover: interpretable convolutional neural networks for improved splice site prediction

AdvanceSplice: Integrating N-gram one-hot encoding and ensemble modeling for enhanced accuracy

DeepSite: bidirectional LSTM and CNN models for predicting DNA–protein binding

Splice site recognition - deciphering Exon-Intron transitions for genetic insights using Enhanced integrated Block-Level gated LSTM model

Splam: a deep-learning-based splice site predictor that improves spliced alignments

DeepSplice: a deep learning approach for accurate prediction of alternative splicing events in the human genome

Predicting splicing patterns from the transcription factor binding sites in the promoter with deep learning

A New Method for Splice Site Prediction Based on the Sequence Patterns of Splicing Signals and Regulatory Elements

CADD-Splice—improving genome-wide variant effect prediction using deep learning-derived splice scores

COSSMO: predicting competitive alternative splice site selection using deep learning

Splice Sites Detection Using Chaos Game Representation and Neural Network.

Deep Motif Dashboard: Visualizing and Understanding Genomic Sequences Using Deep Neural Networks

Reference-informed prediction of alternative splicing and splicing-altering mutations from sequences

SpliceCombo: A Hybrid Technique efficiently use for Principal Component Analysis of Splice Site Prediction

Splicing sites prediction of human genome using machine learning techniques