Abstract:Next-generation RNA sequencing (RNA-seq) technology has been widely used to assess full-length RNA isoform abundance in a high-throughput manner. RNA-seq data offer insight into gene expression levels and transcriptome structures, enabling us to better understand the regulation of gene expression and fundamental biological processes. Accurate isoform quantification from RNA-seq data is challenging due to the information loss in sequencing experiments. A recent accumulation of multiple RNA-seq data sets from the same tissue or cell type provides new opportunities to improve the accuracy of isoform quantification. However, existing statistical or computational methods for multiple RNA-seq samples either pool the samples into one sample or assign equal weights to the samples when estimating isoform abundance. These methods ignore the possible heterogeneity in the quality of different samples and could result in biased and unrobust estimates. In this article, we develop a method, which we call "joint modeling of multiple RNA-seq samples for accurate isoform quantification" (MSIQ), for more accurate and robust isoform quantification by integrating multiple RNA-seq samples under a Bayesian framework. Our method aims to (1) identify a consistent group of samples with homogeneous quality and (2) improve isoform quantification accuracy by jointly modeling multiple RNA-seq samples by allowing for higher weights on the consistent group. We show that MSIQ provides a consistent estimator of isoform abundance, and we demonstrate the accuracy and effectiveness of MSIQ compared with alternative methods through simulation studies on <i>D. melanogaster</i> genes. We justify MSIQ's advantages over existing approaches via application studies on real RNA-seq data from human embryonic stem cells, brain tissues, and the HepG2 immortalized cell line. We also perform a comprehensive analysis of how the isoform quantification accuracy would be affected by RNA-seq sample heterogeneity and different experimental protocols.

Statistical Modeling of RNA-Seq Data

Statistical Inferences for Isoform Expression in RNA-Seq

Modeling and analysis of RNA-seq data: a review from a statistical perspective

Modeling non-uniformity in short-read rates in RNA-Seq data

Joint Estimation of Isoform Expression and Isoform-Specific Read Distribution Using Multisample RNA-Seq Data.

Modeling RNA Degradation for RNA-Seq with Applications

Estimation of Isoform Expression in Rna-Seq Data Using A Hierarchical Bayesian Model

Robust estimation of isoform expression with RNA-Seq data

PennSeq: Accurate Isoform-Specific Gene Expression Quantification in RNA-Seq by Modeling Non-Uniform Read Distribution

A penalized likelihood approach for robust estimation of isoform expression

MSIQ: Joint modeling of multiple RNA-seq samples for accurate isoform quantification

Accurate isoform quantification by joint short- and long-read RNA-sequencing

Statistical modeling of isoform splicing dynamics from RNA-seq time series data

Using non-uniform read distribution models to improve isoform expression inference in RNA-Seq.

Isoform Abundance Inference Provides a More Accurate Estimation of Gene Expression Levels in RNA-seq.

A mixed model approach for joint genetic analysis of alternatively spliced transcript isoforms using RNA-Seq data

Models for transcript quantification from RNA-Seq

A Unified Statistical Framework for Single Cell and Bulk RNA Sequencing Data

RSQ: a Statistical Method for Quantification of Isoform-Specific Structurome Using Transcriptome-Wide Structural Profiling Data

Differential analyses for RNA-seq: transcript-level estimates improve gene-level inferences

A Unified Probabilistic Modeling Framework for Eukaryotic Transcription Based on Nascent RNA Sequencing Data