Abstract:Hyperspectral infrared atmospheric sounding data, characterized by their high vertical resolution, play a crucial role in capturing three-dimensional atmospheric spatial information. The hyperspectral infrared atmospheric detectors HIRAS/HIRAS-II, mounted on the FY3D/EF satellite, have established an initial global coverage network for atmospheric sounding. The collaborative observation approach involving multiple satellites will improve both the coverage and responsiveness of data acquisition, thereby enhancing the overall quality and reliability of the data. In response to the increasing number of channels, the rapid growth of data volume, and the specific requirements of multi-satellite joint observation applications with infrared hyperspectral sounding data, this paper introduces an efficient storage and indexing method for infrared hyperspectral sounding data within a distributed architecture for the first time. The proposed approach, built on the Kubernetes cloud platform, utilizes the Google S2 discrete grid spatial indexing algorithm to establish a grid-based hierarchical model for unified metadata-embedded documents. Additionally, it optimizes the rowkey design using the BPDS model, thereby enabling the distributed storage of data in HBase. The experimental results demonstrate that the query efficiency of the Google S2 grid-based embedded document model is superior to that of the traditional flat model, achieving a query time that is only 35.6% of the latter for a dataset of 5 million records. Additionally, this method exhibits better data distribution characteristics within the global grid compared to the H3 algorithm. Leveraging the BPDS model, the HBase distributed storage system adeptly balances the node load and counteracts the detrimental effects caused by the accumulation of time-series remote sensing images. This architecture significantly enhances both storage and query efficiency, thus laying a robust foundation for forthcoming distributed computing.

Adaptive Indexing for Distributed Array Processing

Distributed High-Dimension Matrix Operation Optimization on Spark

An Efficient and Compact Indexing Scheme for Large-Scale Data Store.

Efficient B-tree Based Indexing for Cloud Data Processing.

Partial Adaptive Indexing for Approximate Query Answering

SciHive: Array-Based Query Processing with HiveQL

Towards Zero-Overhead Adaptive Indexing in Hadoop

In-Memory Indexed Caching for Distributed Data Processing

Progressive online aggregation in a distributed stream system

Query optimization for massively parallel data processing.

Indexing multi-dimensional data in a cloud system.

E3: an Elastic Execution Engine for Scalable Data Processing.

Distributed data management using MapReduce

Using Bitmap Index to Accelerate Accessing Large Scale Scientific Data on Demand

Adaptive Hybrid Indexes

Vhadoop: A Scalable Hadoop Virtual Cluster Platform for MapReduce-Based Parallel Machine Learning with Performance Consideration

Distribution-Based Approach for Efficient Storage and Indexing of Massive Infrared Hyperspectral Sounding Data

A Survey of Spatio-Temporal Big Data Indexing Methods in Distributed Environment

Analytic Queries over Geospatial Time-Series Data Using Distributed Hash Tables

A Framework for Supporting DBMS-like Indexes in the Cloud

Main Memory Adaptive Indexing for Multi-core Systems