Abstract:LLMs with superior response quality--particularly larger or closed-source models--often come with higher inference costs, making their deployment inefficient and costly. Meanwhile, developing foundational LLMs from scratch is becoming increasingly resource-intensive and impractical for many applications. To address the challenge of balancing quality and cost, we introduce Routoo, an architecture designed to optimize the selection of LLMs for specific prompts based on performance, cost, and efficiency. Routoo provides controllability over the trade-off between inference cost and quality, enabling significant reductions in inference costs for a given quality requirement. Routoo comprises two key components: a performance predictor and cost-aware selector. The performance predictor is a lightweight LLM that estimates the expected performance of various underlying LLMs on a given prompt without executing them. The cost-aware selector module then selects the most suitable model based on these predictions and constraints such as cost and latency, significantly reducing inference costs for the same quality. We evaluated Routoo using the MMLU benchmark across 57 domains employing open-source models. Our results show that Routoo matches the performance of the Mixtral 8x7b model while reducing inference costs by one-third. Additionally, by allowing increased costs, Routoo surpasses Mixtral's accuracy by over 5% at equivalent costs, achieving an accuracy of 75.9%. When integrating GPT4 into our model pool, Routoo nearly matches GPT4's performance at half the cost and exceeds it with a 25% cost reduction. These outcomes highlight Routoo's potential to significantly reduce inference costs without compromising quality, and even to establish new state-of-the-art results by leveraging the collective capabilities of multiple LLMs.

Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Efficient Hybrid Inference for LLMs: Reward-Based Token Modelling with Selective Cloud Assistance

RouteLLM: Learning to Route LLMs with Preference Data

PolyRouter: A Multi-LLM Querying System

OptLLM: Optimal Assignment of Queries to Large Language Models

SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization

Performance Characterization of Expert Router for Scalable LLM Inference

A Unified Approach to Routing and Cascading for LLMs

TensorOpera Router: A Multi-Model Router for Efficient LLM Inference

SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models

Efficient and Economic Large Language Model Inference with Attention Offloading

ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency

MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads

On Optimal Caching and Model Multiplexing for Large Model Inference

Hybrid SLM and LLM for Edge-Cloud Collaborative Inference

LLMProxy: Reducing Cost to Access Large Language Models

Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models

One QuantLLM for ALL: Fine-tuning Quantized LLMs Once for Efficient Deployments

RouterBench: A Benchmark for Multi-LLM Routing System

Routoo: Learning to Route to Large Language Models Effectively