Abstract:To deploy machine learning models on-device, practitioners use compression algorithms to shrink and speed up models while maintaining their high-quality output. A critical aspect of compression in practice is model comparison, including tracking many compression experiments, identifying subtle changes in model behavior, and negotiating complex accuracy-efficiency trade-offs. However, existing compression tools poorly support comparison, leading to tedious and, sometimes, incomplete analyses spread across disjoint tools. To support real-world comparative workflows, we develop an interactive visual system called Compress and Compare. Within a single interface, Compress and Compare surfaces promising compression strategies by visualizing provenance relationships between compressed models and reveals compression-induced behavior changes by comparing models' predictions, weights, and activations. We demonstrate how Compress and Compare supports common compression analysis tasks through two case studies, debugging failed compression on generative language models and identifying compression artifacts in image classification models. We further evaluate Compress and Compare in a user study with eight compression experts, illustrating its potential to provide structure to compression workflows, help practitioners build intuition about compression, and encourage thorough analysis of compression's effect on model behavior. Through these evaluations, we identify compression-specific challenges that future visual analytics tools should consider and Compress and Compare visualizations that may generalize to broader model comparison tasks.

Evaluation Discrepancy Discovery: A Sentence Compression Case-study

Toward Human-Like Evaluation for Natural Language Generation with Error Analysis

Accuracy is Not All You Need

With Measured Words: Simple Sentence Selection for Black-Box Optimization of Sentence Compression Algorithms

Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression Experiments

Unsupervised Sentence Compression using Denoising Auto-Encoders

Computational Sentence-level Metrics Predicting Human Sentence Comprehension

OpinSummEval: Revisiting Automated Evaluation for Opinion Summarization

Evaluating Large Language Models for Generalization and Robustness via Data Compression

Re-evaluating Evaluation in Text Summarization

The price of debiasing automatic metrics in natural language evaluation

Investigating a Benchmark for Training-set free Evaluation of Linguistic Capabilities in Machine Reading Comprehension

Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics

Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics

How to Evaluate a Summarizer: Study Design and Statistical Analysis for Manual Linguistic Quality Evaluation

CASPR: Automated Evaluation Metric for Contrastive Summarization

Human Evaluation of Conversations is an Open Problem: comparing the sensitivity of various methods for evaluating dialogue agents

Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation

Learning to Evaluate Image Captioning

Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation

Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies