Abstract:Many LLM tasks are performed in large batches or even offline, and the performance indictor for which is throughput. These tasks usually show the characteristic of prefix sharing, where different prompt input can partially show the common prefix. However, the existing LLM inference engines tend to optimize the streaming requests and show limitations of supporting the large batched tasks with the prefix sharing characteristic. The existing solutions use the LRU-based cache to reuse the KV context of common prefix. The KV context that is about to be reused may prematurely be evicted with the implicit cache management. Even if not evicted, the lifetime of the shared KV context is extended since requests sharing the same context are not scheduled together, resulting in larger memory usage. These streaming oriented systems schedule the requests in the first-come-first-serve or similar order. As a result, the requests with larger ratio of decoding steps may be scheduled too late to be able to mix with the prefill chunks to increase the hardware utilization. Besides, the token and request number based batching can limit the size of token-batch, which keeps the GPU from saturating for the iterations dominated by decoding tokens. We propose BatchLLM to address the above problems. BatchLLM explicitly identifies the common prefixes globally. The requests sharing the same prefix will be scheduled together to reuse the KV context the best, which also shrinks the lifetime of common KV memory. BatchLLM reorders the requests and schedules the requests with larger ratio of decoding first to better mix the decoding tokens with the latter prefill chunks and applies memory-centric token batching to enlarge the token-batch sizes, which helps to increase the GPU utilization. Extensive evaluation shows that BatchLLM outperforms vLLM by 1.1x to 2x on a set of microbenchmarks and two typical industry workloads.

Accelerating NMT Batched Beam Decoding with LMBR Posteriors for Deployment

Lossless Acceleration of Large Language Model via Adaptive N-gram Parallel Decoding

Language-Informed Beam Search Decoding for Multilingual Machine Translation

Beam Search Strategies for Neural Machine Translation

Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding

Chunk-Based Bi-Scale Decoder for Neural Machine Translation.

Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference

Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Later-stage Minimum Bayes-Risk Decoding for Neural Machine Translation

SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Asynchronous and Segmented Bidirectional Encoding for NMT

BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching

Tandem Transformers for Inference Efficient LLMs

Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

BiTA: Bi-Directional Tuning for Lossless Acceleration in Large Language Models

Accelerating Transformer Inference for Translation via Parallel Decoding

Dynamic-Width Speculative Beam Decoding for Efficient LLM Inference

Bi-Decoder Augmented Network for Neural Machine Translation.

The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation

Accelerating LLM Inference with Staged Speculative Decoding

Decoding with Value Networks for Neural Machine Translation.