A compressive seeding algorithm in conjunction with reordering-based compression
Fahu Ji,Qian Zhou,Jue Ruan,Zexuan Zhu,Xianming Liu
DOI: https://doi.org/10.1093/bioinformatics/btae100
IF: 5.8
2024-02-20
Bioinformatics
Abstract:Abstract Motivation Seeding is a rate-limiting stage in sequence alignment for next-generation sequencing (NGS) reads. The existing optimization algorithms typically utilize hardware and machine-learning techniques to accelerate seeding. However, an efficient solution provided by professional NGS compressors has been largely overlooked by far. In addition to achieving remarkable compression ratios by reordering reads, these compressors provide valuable insights for downstream alignment that reveal the repetitive computations accounting for more than 50% of BWA-MEM seeding procedure at common sequencing coverage. Nevertheless, the exploited redundancy information is not fully realized or utilized. Results In this study, we present a compressive seeding algorithm, named CompSeed, to fill the gap. CompSeed, in collaboration with the existing reordering-based compression tools, finishes the BWA-MEM seeding process in about half the time by caching all intermediate seeding results in compact trie structures to directly answer repetitive inquiries that frequently cause random memory accesses. Furthermore, CompSeed demonstrates better performance as sequencing coverage increases, as it focuses solely on the small informative portion of sequencing reads after compression. The innovative strategy highlights the promising potential of integrating sequence compression and alignment to tackle the ever-growing volume of sequencing data. Availability CompSeed is available at https://github.com/i-xiaohu/CompSeed Supplementary information Supplementary data are available at Bioinformatics online.
biochemical research methods,biotechnology & applied microbiology,mathematical & computational biology