Abstract:Data-free knowledge distillation is able to utilize the knowledge learned by a large teacher network to augment the training of a smaller student network without accessing the original training data, avoiding privacy, security, and proprietary risks in real applications. In this line of research, existing methods typically follow an inversion-and-distillation paradigm in which a generative adversarial network on-the-fly trained with the guidance of the pre-trained teacher network is used to synthesize a large-scale sample set for knowledge distillation. In this paper, we reexamine this common data-free knowledge distillation paradigm, showing that there is considerable room to improve the overall training efficiency through a lens of ``small-scale inverted data for knowledge distillation". In light of three empirical observations indicating the importance of how to balance class distributions in terms of synthetic sample diversity and difficulty during both data inversion and distillation processes, we propose Small Scale Data-free Knowledge Distillation SSD-KD. In formulation, SSD-KD introduces a modulating function to balance synthetic samples and a priority sampling function to select proper samples, facilitated by a dynamic replay buffer and a reinforcement learning strategy. As a result, SSD-KD can perform distillation training conditioned on an extremely small scale of synthetic samples (e.g., 10X less than the original training data scale), making the overall training efficiency one or two orders of magnitude faster than many mainstream methods while retaining superior or competitive model performance, as demonstrated on popular image classification and semantic segmentation benchmarks. The code is available at <a class="link-external link-https" href="https://github.com/OSVAI/SSD-KD" rel="external noopener nofollow">this https URL</a>.

Data Distillation: Towards Omni-Supervised Learning

Data-Free Adversarial Distillation

Omni-supervised Facial Expression Recognition via Distilled Data

Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data

Data Distillation: A Survey

Data-to-Model Distillation: Data-Efficient Learning Framework

Low-Resolution Visual Recognition Via Deep Feature Distillation

Learning Efficient Detector with Semi-supervised Adaptive Distillation

Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones

Knowledge Distillation Meets Self-Supervision

Learning Lightweight Object Detectors via Multi-Teacher Progressive Distillation

Small Scale Data-Free Knowledge Distillation

DCD: Discriminative and Consistent Representation Distillation

Scale-Equivalent Distillation for Semi-Supervised Object Detection

A Comprehensive Survey of Dataset Distillation

Self Supervision to Distillation for Long-Tailed Visual Recognition

Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection

Enhancing Dataset Distillation via Label Inconsistency Elimination and Learning Pattern Refinement

Image-to-Lidar Relational Distillation for Autonomous Driving Data

Teaching What You Should Teach: A Data-Based Distillation Method

StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-based 3D Object Detection.