Abstract:Transformer-based Large Language Models (LLMs) have been applied in diverse areas such as knowledge bases, human interfaces, and dynamic agents, and marking a stride towards achieving Artificial General Intelligence (AGI). However, current LLMs are predominantly pretrained on short text snippets, which compromises their effectiveness in processing the long-context prompts that are frequently encountered in practical scenarios. This article offers a comprehensive survey of the recent advancement in Transformer-based LLM architectures aimed at enhancing the long-context capabilities of LLMs throughout the entire model lifecycle, from pre-training through to inference. We first delineate and analyze the problems of handling long-context input and output with the current Transformer-based models. We then provide a taxonomy and the landscape of upgrades on Transformer architecture to solve these problems. Afterwards, we provide an investigation on wildly used evaluation necessities tailored for long-context LLMs, including datasets, metrics, and baseline models, as well as optimization toolkits such as libraries, frameworks, and compilers to boost the efficacy of LLMs across different stages in runtime. Finally, we discuss the challenges and potential avenues for future research. A curated repository of relevant literature, continuously updated, is available at <a class="link-external link-https" href="https://github.com/Strivin0311/long-llms-learning" rel="external noopener nofollow">this https URL</a>.

Domain-specific Chinese Transformer-XL Language Model with Part-of-speech Information

Adapt-and-Distill: Developing Small, Fast and Effective Pretrained Language Models for Domains.

Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Transformer-xl: Language modeling with longer-term dependency

Reinforcing Language Model For Speech Translation With Auxiliary Data

Research on Modeling Units of Transformer Transducer for Mandarin Speech Recognition

Improving Domain Adaptation through Extended-Text Reading Comprehension

Domain-Aware Word Segmentation for Chinese Language: A Document-Level Context-Aware Model

A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

A hybrid Transformer approach for Chinese NER with features augmentation

Adversarial Domain Adaptation For Chinese Semantic Dependency Graph Parsing

Application of the transformer model algorithm in chinese word sense disambiguation: a case study in chinese language

WuDaoCorpora: A Super Large-Scale Chinese Corpora for Pre-Training Language Models

TRAMS: Training-free Memory Selection for Long-range Language Modeling

Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey

A Word Language Model Based Contextual Language Processing On Chinese Character Recognition

Adapting Large Language Models to Domains via Reading Comprehension

Improving Transformer-based Conversational ASR by Inter-Sentential Attention Mechanism

An Improved Chinese Named Entity Recognition Method with TB-LSTM-CRF

Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition

Improving Multi-Party Dialogue Discourse Parsing via Domain Integration