Abstract:This research combines Knowledge Distillation (KD) and Mixture of Experts (MoE) to develop modular, efficient multilingual language models. Key objectives include evaluating adaptive versus fixed alpha methods in KD and comparing modular MoE architectures for handling multi-domain inputs and preventing catastrophic forgetting. KD compresses large language models (LLMs) into smaller, efficient models, while MoE enhances modularity with specialized tasks. Experiments showed similar performance for both KD methods, with marginal improvements from adaptive alpha. A combined loss approach provided more stable learning. The router, trained to classify input sequences into English, French, German, or Python, achieved 99.95% precision, recall, and F1 score, with Logistic Regression being the most effective classifier. Evaluations of modular MoE architectures revealed that Pre-trained Language Experts (PLE) and Joint Expert Embedding Training (JEET) performed similarly, while the MoE with Common Expert (MoE-CE) setup showed slightly lower performance. Including a common expert in MoE-CE improved its performance. Studies on catastrophic forgetting indicated that sequential training led to significant forgetting, while single-session training with balanced batches and the MoE approach mitigated this issue. The MoE architecture preserved knowledge across multiple languages effectively. The research contributes open-sourced resources including the dataset (<a class="link-external link-https" href="https://zenodo.org/doi/10.5281/zenodo.12677631" rel="external noopener nofollow">this https URL</a>), a balanced dataset creation tool (<a class="link-external link-https" href="https://github.com/padas-lab-de/multi-language-dataset-creator" rel="external noopener nofollow">this https URL</a>), and the research codebase (<a class="link-external link-https" href="https://github.com/ModMaamari/mixture-modular-experts" rel="external noopener nofollow">this https URL</a>).

PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning

EvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse Gate

Layerwise Recurrent Router for Mixture-of-Experts

MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing

Mixture of A Million Experts

LEMoE: Advanced Mixture of Experts Adaptor for Lifelong Model Editing of Large Language Models

Mixture of Diverse Size Experts

Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning

MC-MoE: Mixture Compressor for Mixture-of-Experts LLMs Gains More

StableMoE: Stable Routing Strategy for Mixture of Experts

Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design

LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training

Monet: Mixture of Monosemantic Experts for Transformers

UOE: Unlearning One Expert Is Enough For Mixture-of-experts LLMS

Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models

SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget

DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models

Theory on Mixture-of-Experts in Continual Learning

Mixture of Modular Experts: Distilling Knowledge from a Multilingual Teacher into Specialized Modular Language Models

Memory Augmented Language Models through Mixture of Word Experts