Abstract:Data-free knowledge distillation is able to utilize the knowledge learned by a large teacher network to augment the training of a smaller student network without accessing the original training data, avoiding privacy, security, and proprietary risks in real applications. In this line of research, existing methods typically follow an inversion-and-distillation paradigm in which a generative adversarial network on-the-fly trained with the guidance of the pre-trained teacher network is used to synthesize a large-scale sample set for knowledge distillation. In this paper, we reexamine this common data-free knowledge distillation paradigm, showing that there is considerable room to improve the overall training efficiency through a lens of ``small-scale inverted data for knowledge distillation". In light of three empirical observations indicating the importance of how to balance class distributions in terms of synthetic sample diversity and difficulty during both data inversion and distillation processes, we propose Small Scale Data-free Knowledge Distillation SSD-KD. In formulation, SSD-KD introduces a modulating function to balance synthetic samples and a priority sampling function to select proper samples, facilitated by a dynamic replay buffer and a reinforcement learning strategy. As a result, SSD-KD can perform distillation training conditioned on an extremely small scale of synthetic samples (e.g., 10X less than the original training data scale), making the overall training efficiency one or two orders of magnitude faster than many mainstream methods while retaining superior or competitive model performance, as demonstrated on popular image classification and semantic segmentation benchmarks. The code is available at <a class="link-external link-https" href="https://github.com/OSVAI/SSD-KD" rel="external noopener nofollow">this https URL</a>.

Class similarity weighted knowledge distillation for few shot incremental learning

DCCD: Reducing Neural Network Redundancy Via Distillation

Uncertainty-Guided Semi-Supervised Few-Shot Class-Incremental Learning With Knowledge Distillation

Progressive Network Grafting for Few-Shot Knowledge Distillation

Few-Shot Class-Incremental Learning Via Class-Aware Bilateral Distillation

Self-supervised Knowledge Distillation for Few-shot Learning

M2KD: Multi-model and Multi-level Knowledge Distillation for Incremental Learning

Enhancing Few-Shot Learning in Lightweight Models via Dual-Faceted Knowledge Distillation

Few-Shot Incremental Learning for Label-to-Image Translation

Class incremental learning of remote sensing images based on class similarity distillation

Multi-Teacher Knowledge Distillation for Incremental Implicitly-Refined Classification

Incremental Scene Classification Using Dual Knowledge Distillation and Classifier Discrepancy on Natural and Remote Sensing Images

Maintaining Discrimination and Fairness in Class Incremental Learning

Class-Incremental Learning for Remote Sensing Images Based on Knowledge Distillation

Few-Shot Class-Incremental Learning with Non-IID Decentralized Data

Class Incremental Learning with Deep Contrastive Learning and Attention Distillation

An Embarrassingly Simple Approach for Knowledge Distillation

Uncertainty-Aware Contrastive Distillation for Incremental Semantic Segmentation

Hyper-feature aggregation and relaxed distillation for class incremental learning

Channel Distillation: Channel-Wise Attention for Knowledge Distillation

Small Scale Data-Free Knowledge Distillation