Abstract:Compared to traditional learning from scratch, knowledge distillation sometimes makes the DNN achieve superior performance. In this paper, we provide a new perspective to explain the success of knowledge distillation based on the information theory, i.e., quantifying knowledge points encoded in intermediate layers of a DNN for classification. To this end, we consider the signal processing in a DNN as a layer-wise process of discarding information. A knowledge point is referred to as an input unit, the information of which is discarded much less than that of other input units. Thus, we propose three hypotheses for knowledge distillation based on the quantification of knowledge points. 1. The DNN learning from knowledge distillation encodes more knowledge points than the DNN learning from scratch. 2. Knowledge distillation makes the DNN more likely to learn different knowledge points simultaneously. In comparison, the DNN learning from scratch tends to encode various knowledge points sequentially. 3. The DNN learning from knowledge distillation is often more stably optimized than the DNN learning from scratch. To verify the above hypotheses, we design three types of metrics with annotations of foreground objects to analyze feature representations of the DNN, i.e., the quantity and the quality of knowledge points, the learning speed of different knowledge points, and the stability of optimization directions. In experiments, we diagnosed various DNNs on different classification tasks, including image classification, 3D point cloud classification, binary sentiment classification, and question answering, which verified the above hypotheses.

A Closer Look at Knowledge Distillation with Features, Logits, and Gradients

Harmonizing knowledge Transfer in Neural Network with Unified Distillation

Knowledge Distillation Performs Partial Variance Reduction

A Selective Survey on Versatile Knowledge Distillation Paradigm for Neural Network Models

Practical Insights into Knowledge Distillation for Pre-Trained Models

Categories of Response-Based, Feature-Based, and Relation-Based Knowledge Distillation

Class-aware Information for Logit-based Knowledge Distillation

Boosting Knowledge Distillation Via Intra-class Logit Distribution Smoothing

What Knowledge Gets Distilled in Knowledge Distillation?

Lipschitz Continuity Guided Knowledge Distillation

Attention and feature transfer based knowledge distillation

Quantifying the Knowledge in a DNN to Explain Knowledge Distillation for Classification

An Embarrassingly Simple Approach for Knowledge Distillation

Decoupling Dark Knowledge via Block-wise Logit Distillation for Feature-level Alignment

Comparative Knowledge Distillation

Revisiting Knowledge Distillation: an Inheritance and Exploration Framework

What is Lost in Knowledge Distillation?

Distilling Knowledge by Mimicking Features

Knowledge Distillation Via Channel Correlation Structure

Respecting Transfer Gap in Knowledge Distillation

Adaptive Explicit Knowledge Transfer for Knowledge Distillation