Abstract:In this paper, we present a simple yet effective contrastive knowledge distillation approach, which can be formulated as a sample-wise alignment problem with intra- and inter-sample constraints. Unlike traditional knowledge distillation methods that concentrate on maximizing feature similarities or preserving class-wise semantic correlations between teacher and student features, our method attempts to recover the "dark knowledge" by aligning sample-wise teacher and student logits. Specifically, our method first minimizes logit differences within the same sample by considering their numerical values, thus preserving intra-sample similarities. Next, we bridge semantic disparities by leveraging dissimilarities across different samples. Note that constraints on intra-sample similarities and inter-sample dissimilarities can be efficiently and effectively reformulated into a contrastive learning framework with newly designed positive and negative pairs. The positive pair consists of the teacher's and student's logits derived from an identical sample, while the negative pairs are formed by using logits from different samples. With this formulation, our method benefits from the simplicity and efficiency of contrastive learning through the optimization of InfoNCE, yielding a run-time complexity that is far less than $O(n^2)$, where $n$ represents the total number of training samples. Furthermore, our method can eliminate the need for hyperparameter tuning, particularly related to temperature parameters and large batch sizes. We conduct comprehensive experiments on three datasets including CIFAR-100, ImageNet-1K, and MS COCO. Experimental results clearly confirm the effectiveness of the proposed method on both image classification and object detection tasks. Our source codes will be publicly available at

Classifier Reuse-Based Contrastive Knowledge Distillation

DCCD: Reducing Neural Network Redundancy Via Distillation

Category contrastive distillation with self-supervised classification

Knowledge Distillation with the Reused Teacher Classifier

ADCL: Adversarial Distilled Contrastive Learning on lightweight models for self-supervised image classification

Knowledge Distillation Meets Self-Supervision

Knowledge distillation based on projector integration and classifier sharing

Knowledge Distillation with Deep Supervision

Hybrid mix-up contrastive knowledge distillation

CKD: Contrastive Knowledge Distillation from A Sample-wise Perspective

DCD: Discriminative and Consistent Representation Distillation

Teacher-Student Complementary Sample Contrastive Distillation

DistilCSE: Effective Knowledge Distillation For Contrastive Sentence Embeddings

Knowledge Distillation with a Precise Teacher and Prediction with Abstention

Contrastive Representation Distillation

Multi-teacher Contrastive Knowledge Inversion for Data-Free Distillation

Learning Lightweight Object Detectors via Multi-Teacher Progressive Distillation

Similarity Transfer for Knowledge Distillation

Distilling Knowledge via Intermediate Classifiers

Distilling Image Classifiers in Object Detectors

Multi-instance semantic similarity transferring for knowledge distillation