Abstract:Pre-trained language models (PLMs) like BERT have made great progress in NLP. News articles usually contain rich textual information, and PLMs have the potentials to enhance news text modeling for various intelligent news applications like news recommendation and retrieval. However, most existing PLMs are in huge size with hundreds of millions of parameters. Many online news applications need to serve millions of users with low latency tolerance, which poses huge challenges to incorporating PLMs in these scenarios. Knowledge distillation techniques can compress a large PLM into a much smaller one and meanwhile keeps good performance. However, existing language models are pre-trained and distilled on general corpus like Wikipedia, which has some gaps with the news domain and may be suboptimal for news intelligence. In this paper, we propose NewsBERT, which can distill PLMs for efficient and effective news intelligence. In our approach, we design a teacher-student joint learning and distillation framework to collaboratively learn both teacher and student models, where the student model can learn from the learning experience of the teacher model. In addition, we propose a momentum distillation method by incorporating the gradients of teacher model into the update of student model to better transfer useful knowledge learned by the teacher model. Extensive experiments on two real-world datasets with three tasks show that NewsBERT can effectively improve the model performance in various intelligent news applications with much smaller models.

Comparative analysis of strategies of knowledge distillation on BERT for text matching

Adapt-and-Distill: Developing Small, Fast and Effective Pretrained Language Models for Domains.

LAD: Layer-Wise Adaptive Distillation for BERT Model Compression

Improving Knowledge Distillation for BERT Models: Loss Functions, Mapping Methods, and Weight Tuning

Towards Non-task-specific Distillation of BERT via Sentence Representation Approximation

Patient Knowledge Distillation for BERT Model Compression

Mandarin Text-to-Speech Front-End with Lightweight Distilled Convolution Network

NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application

An Empirical Study of Uniform-Architecture Knowledge Distillation in Document Ranking

Extremely Small BERT Models from Mixed-Vocabulary Training

Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation

Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation

DDK: Dynamic structure pruning based on differentiable search and recursive knowledge distillation for BERT

XtremeDistil: Multi-stage Distillation for Massive Multilingual Models

One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers

Knowledge distillation and data augmentation for NLP light pre-trained models

Which Student is Best? A Comprehensive Knowledge Distillation Exam for Task-Specific BERT Models

A Comparative Analysis of Task-Agnostic Distillation Methods for Compressing Transformer Language Models

MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models

Research of Weibo Text Classification Based on Knowledge Distillation and Joint Model

Distilling BERT for low complexity network training