ANASFV: a workflow for ASFV whole genome sequencing, assembly, and evaluation

Ke Li,Xu Han,Yanwen Shao,Yiwen Zhang,Xiaomin Zhao,Eager Wayne Johnson,Runsheng Li
DOI: https://doi.org/10.1101/2024.07.08.602471
2024-09-18
Abstract:African Swine Fever Virus (ASFV) poses an ongoing threat with widespread outbreaks affecting both domestic and wild pig populations globally. Effective management of ASFV outbreaks necessitates a deep understanding of its genetic diversity and evolutionary dynamics. Despite the advantages of nanopore sequencing for genome analysis, its application to ASFV genomes encounters specific challenges, such as high host DNA contamination that reduces viral genome coverage and an inherently higher error rate producing small insertions and deletions (indels). Another notable issue is the lack of standardized methods for assessing the quality of ASFV genomes. Furthermore, an increasing number of recombinant isolates of genotype I and II, have been observed, further complicating the task of phylogenetic analysis. To overcome these obstacles, we developed the ANASFV (analysis of an ASFV) pipeline. The pipeline is used to solve four tasks: In the first task, the pipeline introduces an amplicon sequencing approach that significantly improves genomic coverage, enabling reliable genome assembly. The pipeline also introduce reference-aided polishing techniques to correct small indels caused by nanopore sequencing errors. Moving on to the second task, a genome quality assessment system was established to evaluate the completeness and accuracy of the assembled genomes. For the third task, a method to rapidly analyze whether an isolate is a recombinant between genotype I and II, as well as to determine the pattern of recombination, based on gene similarity. Lastly, a comprehensive phylogenetic analysis based on coding sequences (CDS) was conducted, allowing us to generate a refined phylogenetic tree that includes all known ASFV genomes. The ANASFV pipeline will facilitate ASFV full genome sequencing using the nanopore platform. The pipeline will also support robust downstream bioinformatic analyses to assess genome integrity and conduct detailed phylogenetic studies based on whole-genome data.
Genetics
What problem does this paper attempt to address?