ScaleQC: a scalable lossy to lossless solution for NGS data compression

Rongshan Yu,Wenxian Yang
DOI: https://doi.org/10.1093/bioinformatics/btaa543
IF: 5.8
2020-01-01
Bioinformatics
Abstract:Motivation: Per-base quality values in Next Generation Sequencing data take a significant portion of storage even after compression. Lossy compression technologies could further reduce the space used by quality values. However, in many applications, lossless compression is still desired. Hence, sequencing data in multiple file formats have to be prepared for different applications. Results: We developed a scalable lossy to lossless compression solution for quality values named ScaleQC (Scalable Quality value Compression). ScaleQC is able to provide the so-called bit-stream level scalability that the losslessly compressed bit-stream by ScaleQC can be further truncated to lower data rates without incurring an expensive transcoding operation. Despite its scalability, ScaleQC still achieves comparable compression performance at both lossless and lossy data rates compared to the existing lossless or lossy compressors.
What problem does this paper attempt to address?