Abstract:A study has proposed a massive data storage method for electronic archive management based on HBase, aiming to improve the intelligence, efficiency and retrieval performance of data storage. The results show that this method is superior to traditional database systems in terms of write speed and query latency and is suitable for efficient storage and management of massive electronic archives. The acceleration of the digitalization process in enterprise and university education management has generated a massive amount of electronic archive data. In order to improve the intelligence, storage quality, and efficiency of electronic records management and achieve efficient storage and fast retrieval of data storage models, this study proposes a massive data storage model based on HBase and its retrieval optimization scheme design. In addition, HDFS is introduced to construct a two‐level storage structure and optimize values to improve the scalability and load balancing of HBase, and the retrieval efficiency of the HBase storage model is improved through SL‐TCR and BF filters. The results indicated that HDFS could automatically recover data after node, network partition, and NameNode failures. The write time of HBase was 56 s, which was 132 and 246 s less than Cassandra and CockroachDB. The query latency was reduced by 23% and 32%, and the query time was reduced by 9988.51 ms, demonstrating high reliability and efficiency. The delay of BF‐SL‐TCL was 1379.28 s after 1000 searches, which was 224.78 and 212.74 s less than SL‐TCL and Blockchain Retrieval Acceleration and reduced the delay under high search times. In summary, this storage model has obvious advantages in storing massive amounts of electronic archive data and has high security and retrieval efficiency, which provides important reference for the design of storage models for future electronic archive management. The storage model designed by the research institute has obvious advantages in storing massive electronic archive data, solving the problem of lack of scalability in electronic archive management when facing massive data, and has high security and retrieval efficiency. It has important reference for the design of storage models for future electronic archive management.

Small files access efficiency in hadoop distributed file system a case study performed on British library text files

A digital library architecture supporting massive small files and efficient replica maintenance.

A Novel Scalable Architecture of Cloud Storage System for Small Files Based on P2P

A Proposed Approach for Improving Hadoop Performance for Handling Small Files

An archive‐based method for efficiently handling small file problems in HDFS

CSFC: A New Centroid Based Clustering Method to Improve the Efficiency of Storing and Accessing Small Files in Hadoop

A Novel Approach to Improving the Efficiency of Storing and Accessing Small Files on Hadoop: A Case Study by PowerPoint Files

Small Files Problem Resolution via Hierarchical Clustering Algorithm

RCFile: A Fast and Space-Efficient Data Placement Structure in MapReduce-based Warehouse Systems

Data Management Techniques in Hadoop Framework for Handling Small Files: A Survey

Addressing the Small Files Issue in Hadoop

Impact of Small Files on Hadoop Performance: Literature Survey and Open Points

Pseudo-Cache-Based IoT Small Files Management Framework in HDFS Cluster

Storage-Optimization Method for Massive Small Files of Agricultural Resources Based on Hadoop

Survey on Resource Management Solutions to Speed up Processing Small Files in Hadoop Cluster

Contributions to Hadoop File System Architecture by Revising the File System Usage Along with Automatic Service

A Novel Approach for Improving Security and Storage Efficiency on HDFS

Small File Read Performance Optimization Based on Redis Cache in Distributed Storage System

A Packaging Approach for Massive Amounts of Small Geospatial Files with HDFS.

Elastic HDFS: interconnected distributed architecture for availability–scalability enhancement of large-scale cloud storages

Massive Data HBase Storage Method for Electronic Archive Management