Abstract:Data-free knowledge distillation is able to utilize the knowledge learned by a large teacher network to augment the training of a smaller student network without accessing the original training data, avoiding privacy, security, and proprietary risks in real applications. In this line of research, existing methods typically follow an inversion-and-distillation paradigm in which a generative adversarial network on-the-fly trained with the guidance of the pre-trained teacher network is used to synthesize a large-scale sample set for knowledge distillation. In this paper, we reexamine this common data-free knowledge distillation paradigm, showing that there is considerable room to improve the overall training efficiency through a lens of ``small-scale inverted data for knowledge distillation". In light of three empirical observations indicating the importance of how to balance class distributions in terms of synthetic sample diversity and difficulty during both data inversion and distillation processes, we propose Small Scale Data-free Knowledge Distillation SSD-KD. In formulation, SSD-KD introduces a modulating function to balance synthetic samples and a priority sampling function to select proper samples, facilitated by a dynamic replay buffer and a reinforcement learning strategy. As a result, SSD-KD can perform distillation training conditioned on an extremely small scale of synthetic samples (e.g., 10X less than the original training data scale), making the overall training efficiency one or two orders of magnitude faster than many mainstream methods while retaining superior or competitive model performance, as demonstrated on popular image classification and semantic segmentation benchmarks. The code is available at <a class="link-external link-https" href="https://github.com/OSVAI/SSD-KD" rel="external noopener nofollow">this https URL</a>.

Data-Free Ensemble Knowledge Distillation for Privacy-conscious Multimedia Model Compression.

Model Compression via Collaborative Data-Free Knowledge Distillation for Edge Intelligence.

DCCD: Reducing Neural Network Redundancy Via Distillation

CDFKD-MFS: Collaborative Data-free Knowledge Distillation Via Multi-level Feature Sharing

Data-Free Adversarial Distillation

CDFKD-MFS: Collaborative Data-free Knowledge Distillation via Multi-level Feature Sharing

Dual Discriminator Adversarial Distillation for Data-free Model Compression

A Model Compression Method Using Significant Data and Knowledge Distillation

Up to 100x Faster Data-Free Knowledge Distillation

Exploring Model Compression Limits and Laws: A Pyramid Knowledge Distillation Framework for Satellite-on-Orbit Object Recognition

Ability-aware knowledge distillation for resource-constrained embedded devices

Residual Error Based Knowledge Distillation

Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation

Communication-efficient federated learning via knowledge distillation

FedNKD: A Dependable Federated Learning Using Fine-tuned Random Noise and Knowledge Distillation

Deep Collective Knowledge Distillation

Channel-Correlation-Based Selective Knowledge Distillation

Data Efficient Stagewise Knowledge Distillation

Small Scale Data-Free Knowledge Distillation

Knowledge Distillation as Efficient Pre-training: Faster Convergence, Higher Data-efficiency, and Better Transferability

De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts