Abstract:Continual learning aims to rapidly and continually learn the current task from a sequence of tasks, using the knowledge obtained in the past, while performing well on prior tasks. A key challenge in this setting is the stability–plasticity dilemma existing in current and previous tasks, i.e., a high-stability network is weak to learn new knowledge in an effort to maintain previous knowledge. Correspondingly, a high-plasticity network can easily forget old tasks while dealing with well on the new task. Compared to other kinds of methods, the methods based on experience replay have shown great advantages to overcome catastrophic forgetting. One common limitation of this method is the data imbalance between the previous and current tasks, which would further aggravate forgetting. Moreover, how to effectively address the stability–plasticity dilemma in this setting is also an urgent problem to be solved. In this paper, we overcome these challenges by proposing a novel framework called Meta-learning update via Multi-scale Knowledge Distillation and Data Augmentation (MMKDDA). Specifically, we apply multi-scale knowledge distillation to grasp the evolution of long-range and short-range spatial relationships at different feature levels to alleviate the problem of data imbalance. Besides, our method mixes the samples from the episodic memory and current task in the online continual training procedure, thus alleviating the side influence due to the change of probability distribution. Moreover, we optimize our model via the meta-learning update by resorting to the number of tasks seen previously, which is helpful to keep a better balance between stability and plasticity. Finally, our extensive experiments on four benchmark datasets show the effectiveness of the proposed MMKDDA framework against other popular baselines, and ablation studies are also conducted to further analyze the role of each component in our framework.

Meta-Learning via Feature-Label Memory Network

Labeled Memory Networks for Online Model Adaptation

A Survey of Meta-learning for Classification Tasks

MetaMIML: Meta Multi-Instance Multi-Label Learning

One-shot Learning with Memory-Augmented Neural Networks

Memory transformation networks for weakly supervised visual classification

Any-Way Meta Learning

MCML: A Novel Memory-based Contrastive Meta-Learning Method for Few Shot Slot Tagging

Representation Based and Attention Augmented Meta Learning

Attention-Augmented Memory Network for Image Multi-Label Classification

Decomposed Meta-Learning for Few-Shot Sequence Labeling

Multimodal Meta-Learning for Time Series Regression

Memory and attention in deep learning

Online Continual Learning Via the Meta-learning Update with Multi-scale Knowledge Distillation and Data Augmentation

Concept learning through deep reinforcement learning with memory-augmented neural networks

Meta-LMTC - Meta-Learning for Large-Scale Multi-Label Text Classification.

Multi-Domain Learning by Meta-Learning: Taking Optimal Steps in Multi-Domain Loss Landscapes by Inner-Loop Learning

Meta Feature Modulator for Long-tailed Recognition

A Novel Hierarchical Adaptive Feature Fusion Method for Meta-Learning

MLAAN: Scaling Supervised Local Learning with Multilaminar Leap Augmented Auxiliary Network

Distributed Associative Memory Network with Memory Refreshing Loss