Abstract:Time series play a crucial role in many fields, including finance, healthcare, industry, and environmental monitoring. The storage and retrieval of time series can be challenging due to their unstoppable growth. In fact, these applications often sacrifice precious historical data to make room for new data. General-purpose compressors can mitigate this problem with their good compression ratios, but they lack efficient random access on compressed data, thus preventing real-time analyses. Ad-hoc streaming solutions, instead, typically optimise only for compression and decompression speed, while giving up compression effectiveness and random access functionality. Furthermore, all these methods lack awareness of certain special regularities of time series, whose trends over time can often be described by some linear and nonlinear functions. To address these issues, we introduce NeaTS, a randomly-accessible compression scheme that approximates the time series with a sequence of nonlinear functions of different kinds and shapes, carefully selected and placed by a partitioning algorithm to minimise the space. The approximation residuals are bounded, which allows storing them in little space and thus recovering the original data losslessly, or simply discarding them to obtain a lossy time series representation with maximum error guarantees. Our experiments show that NeaTS improves the compression ratio of the state-of-the-art lossy compressors that use linear or nonlinear functions (or both) by up to 14%. Compared to lossless compressors, NeaTS emerges as the only approach to date providing, simultaneously, compression ratios close to or better than the best existing compressors, a much faster decompression speed, and orders of magnitude more efficient random access, thus enabling the storage and real-time analysis of massive and ever-growing amounts of (historical) time series data.

Elf: Erasing-based Lossless Floating-Point Compression.

Erasing-based lossless compression method for streaming floating-point time series

Adaptive Encoding Strategies for Erasing-Based Lossless Floating-Point Compression

High Performance Lossless Compression of Scientific Floating Data

Lossless preprocessing of floating point data to enhance compression

A Versatile Compression Method for Floating-Point Data Stream

A fast algorithm for DEM lossless compression based on embedded wavelet coding

Fast Algorithm for DEM Lossless Compression

Change a Bit to save Bytes: Compression for Floating Point Time-Series Data

A Lossless Electrocardiogram Compression System Based on Dual-Mode Prediction and Error Modeling.

A Fast Compression Algorithm for Seismic Data from Non-Cable Seismographs

Use cases of lossy compression for floating-point data in scientific data sets

Machete: an Efficient Lossy Floating-Point Compressor Designed for Time Series Databases

Fast Lossless Neural Compression with Integer-Only Discrete Flows

A Universal 4D Model for Double-Efficient Lossless Data Compressions

EFloat: Entropy-coded Floating Point Format for Compressing Vector Embedding Models

Real-Time Lossless Compression for Ultrahigh-Density Synchrophasor and Point-on-Wave Data

A Novel Method of Lossless Compression for 2-D Astronomical Spectra Images

Integer Network for Cross Platform Graph Data Lossless Compression

Learned Compression of Nonlinear Time Series With Random Access

iFlow: Numerically Invertible Flows for Efficient Lossless Compression via a Uniform Coder