Abstract:Recent CNNs (convolutional neural networks) have become more and more compact. The elegant structure design highly improves the performance of CNNs. With the development of knowledge distillation technique, the performance of CNNs gets further improved. However, existing knowledge distillation guided methods either rely on offline pretrained high-quality large teacher models or online heavy training burden. To solve the above problems, we propose a feature-sharing and weight-sharing based ensemble network (training framework) guided by knowledge distillation (EKD-FWSNet) to make baseline models stronger in terms of representation ability with less training computation and memory cost involved. Specifically, motivated by getting rid of the dependence of offline pretrained teacher model, we design an end-to-end online training scheme to optimize EKD-FWSNet. Motivated by decreasing the online training burden, we only introduce one auxiliary classmate branch to construct multiple forward branches, which will then be integrated as ensemble teacher to guide baseline model. Compared to previous online ensemble training frameworks, EKD-FWSNet can provide diverse output predictions without relying on increasing auxiliary classmate branches. Motivated by maximizing the optimization power of EKD-FWSNet, we exploit the representation potential of weight-sharing blocks and design efficient knowledge distillation mechanism in EKD-FWSNet. Extensive comparison experiments and visualization analysis on benchmark datasets (CIFAR-10/100, tiny-ImageNet, CUB-200 and ImageNet) show that self-learned EKD-FWSNet can boost the performance of baseline models by large margin, which has obvious superiority compared to previous related methods. Extensive analysis also proves the interpretability of EKD-FWSNet. Our code is available at https://github.com/cv516Buaa/EKD-FWSNet.

Investigating the Impact of Weight Sharing Decisions on Knowledge Transfer in Continual Learning

Progressive Learning without Forgetting

Achieving Forgetting Prevention and Knowledge Transfer in Continual Learning

Learning After Learning: Positive Backward Transfer in Continual Learning

A Deeper Knowledge Tracking Model Integrating Cognitive Theory and Learning Behavior

Learn by Oneself: Exploiting Weight-Sharing Potential in Knowledge Distillation Guided Ensemble Network

Federated Continual Learning with Weighted Inter-client Transfer

Optimizing Reusable Knowledge for Continual Learning via Metalearning

Imbalance Mitigation for Continual Learning via Knowledge Decoupling and Dual Enhanced Contrastive Learning

An Experimental Survey of Incremental Transfer Learning for Multicenter Collaboration

Beyond Not-Forgetting: Continual Learning with Backward Knowledge Transfer

Continual Learning with Dependency Preserving Hypernetworks

Sub-network Discovery and Soft-masking for Continual Learning of Mixed Tasks

Exclusive Supermask Subnetwork Training for Continual Learning

Continual Learning in the Teacher-Student Setup: Impact of Task Similarity

A Survey of Incremental Transfer Learning: Combining Peer-to-Peer Federated Learning and Domain Incremental Learning for Multicenter Collaboration

Does Continual Learning Equally Forget All Parameters?

Mitigating Interference in the Knowledge Continuum through Attention-Guided Incremental Learning

Deeper Insights into Weight Sharing in Neural Architecture Search

Disentangling and Mitigating the Impact of Task Similarity for Continual Learning

Fantastic Gains and Where to Find Them: On the Existence and Prospect of General Knowledge Transfer between Any Pretrained Model