Abstract:Increasingly, model compression techniques enable large language models (LLMs) to be deployed in real-world applications. As a result of this momentum towards local deployment, compressed LLMs will interact with a large population. Prior work on compression typically prioritize preserving perplexity, which is directly analogous to training loss. The impact of compression method on other critical aspects of model behavior\, -- \,particularly safety\, -- \,requires systematic assessment. To this end, we investigate the impact of model compression along four dimensions: (1) degeneration harm, i.e., bias and toxicity in generation; (2) representational harm, i.e., biases in discriminative tasks; (3) dialect bias; and(4) language modeling and downstream task performance. We examine a wide spectrum of LLM compression techniques, including unstructured pruning, semi-structured pruning, and quantization. Our analysis reveals that compression can lead to unexpected consequences. Although compression may unintentionally alleviate LLMs' degeneration harm, it can still exacerbate representational harm. Furthermore, increasing compression produces a divergent impact on different protected groups. Finally, different compression methods have drastically different safety impacts: for example, quantization mostly preserves bias while pruning degrades quickly. Our findings underscore the importance of integrating safety assessments into the development of compressed LLMs to ensure their reliability across real-world applications.\footnote{Our implementation and results are available here: \url{<a class="link-external link-https" href="https://github.com/zhichaoxu-shufe/Beyond-Perplexity-Compression-Safety-Eval" rel="external noopener nofollow">this https URL</a>}}

Evaluating Zero-Shot Long-Context LLM Compression

Ranking LLMs by compression

Retaining Key Information under High Compression Ratios: Query-Guided Compressor for LLMs

Aggressive Post-Training Compression on Extremely Large Language Models

Compression Represents Intelligence Linearly

Beyond Perplexity: Multi-dimensional Safety Evaluation of LLM Compression

Extending Context Window of Large Language Models via Semantic Compression

Evaluating Large Language Models for Generalization and Robustness via Data Compression

Understanding is Compression

A Survey on Model Compression for Large Language Models

Enhancing and Accelerating Large Language Models via Instruction-Aware Contextual Compression

Evaluating the Impact of Compression Techniques on Task-Specific Performance of Large Language Models

LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit

LLM Vocabulary Compression for Low-Compute Environments

Zero-Delay QKV Compression for Mitigating KV Cache and Network Bottlenecks in LLM Inference

LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment

LLMZip: Lossless Text Compression using Large Language Models

Efficient Large Multi-modal Models via Visual Context Compression

Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression

Training LLMs over Neurally Compressed Text

In-Context Former: Lightning-fast Compressing Context for Large Language Model