CFC-seq: identification of full-length capped RNAs unveil enhancer-derived transcription

Chi Wai Yip,Callum Parr,Hazuki Takahashi,Kayoko Yasuzawa,Matthew Valentine,Hiromi Nishiyori-Sueki,Camilla Ugolini,Valeria Ranzani,Mitsuyoshi Murata,Masaki Kato,Wenjing Kang,Wing Hin Yip,Youtaro Shibayama,Andre Darah Sim,Ying Chen,Xufeng Shu,Jonathan Darah Moody,Ramzan Umarov,Jen-Chien Chang,Luca Pandolfini,Tsugumi Kawashima,Michihira Tagami,Tomoe Nobusada,Tsukasa Kouno,Carlos Alfonso Gonzalez,Roberto Albanese,Francesco Dossena,Nejc Haberman,Kokoro Ozaki,Takeya Kasukawa,Boris Lenhard,Martin Frith,Beatrice Bodega,Francesco Nicassio,Lorenzo Calviello,Magda Bienko,Ivano Legnini,Valerie Hilgers,Stefano Gustincich,Jonathan Goeke,Charles Henri Lecellier,Jay W Shin,Chung-Chau Hon,Piero Carninci
DOI: https://doi.org/10.1101/2024.10.31.620483
2024-11-01
Abstract:Long-read sequencing has emerged as a powerful tool for uncovering novel transcripts and genes. However, existing protocols often lack confidence in identifying the transcription start site (TSS) and fail to capture non-poly(A) RNA, thereby limiting the discovery of novel genes, particularly long non-coding RNAs (lncRNAs). In this study, we introduce Cap-trap full-length cDNA sequencing (CFC-seq), a comprehensive protocol that combines Cap-trapping and poly(A)-tailing with Oxford Nanopore sequencing. This protocol enables precise identification of TSSs and full-length transcripts. Applying CFC-seq to two in vitro differentiation time courses resulted in approximately 236 million mappable reads. The transcript Start-site Aware Long-read Assembler (SALA) was developed for de novo assembling the transcript models, leading to the identification of 39,425 confident novel genes. Using this dataset, enhancer-derived ncRNAs were re-defined with longer length and more splicing activity, which were correlated with enhancer structure. Compared to enhancers with CpG islands, TATA box enhancers were shown to be more cell-type specific with fewer chromatin interaction but produced longer and more stable polyadenylated RNA. A significant proportion of these TATA box-derived eRNAs originated from LTR transposable elements. Overall, this study systematically annotated ~24,000 novel eRNA genes and correlated their transcription properties with enhancer structure.
Genomics
What problem does this paper attempt to address?
This paper aims to address the shortcomings of existing long - read - length sequencing technologies in identifying transcription start sites (TSSs) and capturing non - polyadenylated RNA (non - poly(A) RNA), which limit the discovery of novel genes, especially long non - coding RNAs (lncRNAs). To this end, the researchers developed a new method named Cap - trap full - length cDNA sequencing (CFC - seq), which combines Cap - trapping technology and polyadenylation (poly(A) - tailing) and uses the Oxford Nanopore sequencing platform for long - read - length sequencing. This method can accurately identify TSSs and full - length transcripts, thereby enhancing the ability to discover novel genes, especially those enhancer - derived non - coding RNAs (eRNAs). By applying the CFC - seq technology, the researchers were able to: 1. **Accurately identify transcription start sites (TSSs)**: CFC - seq captures RNA molecules with 5' - end cap structures through Cap - trapping technology, enabling more accurate identification of TSSs, which is particularly important for distinguishing independently transcribed lncRNAs from other genes. 2. **Capture non - polyadenylated RNA**: By adding poly(A) tails to total RNA, CFC - seq can effectively capture non - coding RNAs lacking natural poly(A) tails, which are often excluded in standard long - read - length sequencing. 3. **Discover a large number of novel genes**: The researchers obtained approximately 236 million mappable reads through the application of two in - vitro differentiation time courses, and used the developed SALA assembler to assemble transcript models de novo, ultimately identifying 39,425 novel genes with true TSS support, including approximately 24,000 eRNA genes. 4. **Redefine enhancer - derived non - coding RNAs**: The study found that these eRNAs are generally longer, have more splicing activities, and are related to enhancer structures. Compared with enhancers containing CpG islands, enhancers containing TATA boxes exhibit stronger cell - type specificity but fewer chromatin interactions and produce longer and more stable polyadenylated RNAs. In addition, a considerable proportion of TATA - box - derived eRNAs are derived from long terminal repeat (LTR) transposable elements. In conclusion, the CFC - seq technology systematically annotates approximately 24,000 novel eRNA genes and correlates their transcriptional characteristics with the enhancer structures from which they are derived, providing a new perspective for in - depth understanding of enhancer functions and their regulatory mechanisms.