Improved Rat Genome Gene Prediction by Integration of ESTs with RNA-Seq Information

Liping Li,Enguo Chen,Chun Yang,Jun Zhu,Pushkala Jayaraman,Jeffrey De Pons,Catherine C. Kaczorowski,Howard J. Jacob,Andrew S. Greene,Matthew R. Hodges,Allen W. Cowley,Mingyu Liang,Haiming Xu,Pengyuan Liu,Yan Lu
DOI: https://doi.org/10.1093/bioinformatics/btu608
IF: 5.8
2014-01-01
Bioinformatics
Abstract:MOTIVATIONRNA-Seq (also called whole-transcriptome sequencing) is an emerging technology that uses the capabilities of next-generation sequencing to detect and quantify entire transcripts. One of its important applications is the improvement of existing genome annotations. RNA-Seq provides rapid, comprehensive and cost-effective tools for the discovery of novel genes and transcripts compared with expressed sequence tag (EST), which is instrumental in gene discovery and gene sequence determination. The rat is widely used as a laboratory disease model, but has a less well-annotated genome as compared with humans and mice. In this study, we incorporated deep RNA-Seq data from three rat tissues-bone marrow, brain and kidney-with EST data to improve the annotation of the rat genome.RESULTSOur analysis identified 32 197 transcripts, including 13 461 known transcripts, 13 934 novel isoforms and 4802 new genes, which almost doubled the numbers of transcripts in the current public rat genome database (rn5). Comparisons of our predicted protein-coding gene sets with those in public datasets suggest that RNA-Seq significantly improves genome annotation and identifies novel genes and isoforms in the rat. Importantly, the large majority of novel genes and isoforms are supported by direct evidence of RNA-Seq experiments. These predicted genes were integrated into the Rat Genome Database (RGD) and can serve as an important resource for functional studies in the research community.AVAILABILITY AND IMPLEMENTATIONThe predicted genes are available at http://rgd.mcw.edu.
What problem does this paper attempt to address?