Abstract:As Large Language Models (LLMs) continue to evolve, more are being designed to handle long-context inputs. Despite this advancement, many models face challenges in achieving high precision on long-context tasks, often showing a ``lost in the middle'' issue. This paper identifies the root of these issues as a deficiency in retrieval capabilities, exacerbated by the sparsity of key information in long contexts. To tackle this challenge, we introduce a novel approach called ``Paraphrasing the Original Text'', aimed at augmenting LLMs' proficiency in extracting information from long context. This enhancement is achieved through a specialized supervised fine-tuning stage that incorporates paraphrasing information into training samples, thereby improving the model's retrieval capabilities for long-context scenarios. Testing on datasets like LongBench and NaturalQuestions Multi-document QA dataset, our method demonstrated significant improvements in managing long-context tasks, effectively addressing the ``lost in the middle'' dilemma. Specifically, we observed an average performance increase of 6.4\% and 5.9\% across these datasets, respectively. Moreover, our approach is efficient, requiring minimal overhead with fine-tuning needed on just 19k samples. The model and training data have been made available on HuggingFace(<a class="link-external link-https" href="https://huggingface.co/yuyijiong/Qwen-14b-chat-yarn-32k" rel="external noopener nofollow">this https URL</a>).

An Empirical Study on Context Length for Open-Domain Dialog Generation

How To Make Context More Useful? An Empirical Study On Context-Aware Neural Conversational Models

An Empirical Investigation of Pre-Trained Transformer Language Models for Open-Domain Dialogue Generation

Empower Your Model with Longer and Better Context Comprehension

Building Context-Related Dialogue Systems Based on Chinese-Script-Dialogue Corpus

Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts

How to Represent Context Better? an Empirical Study on Context Modeling for Multi-turn Response Selection.

Extending the Transformer with Context and Multi-dimensional Mechanism for Dialogue Response Generation.

Local and Global Contexts for Conversation

RULER: What's the Real Context Size of Your Long-Context Language Models?

Do Long-Range Language Models Actually Use Long-Range Context?

Exploring Context Window of Large Language Models via Decomposed Positional Vectors

What Kinds of Tokens Benefit from Distant Text? An Analysis on Long Context Language Modeling

Improving Contextual Language Models for Response Retrieval in Multi-Turn Conversation

Reasoning in Dialog: Improving Response Generation by Context Reading Comprehension

Training With "Paraphrasing the Original Text'' Improves Long-Context Performance

Improvement of a dedicated model for open domain persona-aware dialogue generation

Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Unsupervised Context Rewriting for Open Domain Conversation

Improving the Transformer Translation Model with Document-Level Context

X-RECOSA: Multi-Scale Context Aggregation for Multi-Turn Dialogue Generation