RouterBench: A Benchmark for Multi-LLM Routing System

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Qitian Jason, Bieker, Jacob, Li, Xiuyu, Jiang, Nan, Keigwin, Benjamin, Ranganath, Gaurav, Keutzer, Kurt, Upadhyay, Shriyash Kaustubh
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909153781350400
author Hu, Qitian Jason
Bieker, Jacob
Li, Xiuyu
Jiang, Nan
Keigwin, Benjamin
Ranganath, Gaurav
Keutzer, Kurt
Upadhyay, Shriyash Kaustubh
author_facet Hu, Qitian Jason
Bieker, Jacob
Li, Xiuyu
Jiang, Nan
Keigwin, Benjamin
Ranganath, Gaurav
Keutzer, Kurt
Upadhyay, Shriyash Kaustubh
contents As the range of applications for Large Language Models (LLMs) continues to grow, the demand for effective serving solutions becomes increasingly critical. Despite the versatility of LLMs, no single model can optimally address all tasks and applications, particularly when balancing performance with cost. This limitation has led to the development of LLM routing systems, which combine the strengths of various models to overcome the constraints of individual LLMs. Yet, the absence of a standardized benchmark for evaluating the performance of LLM routers hinders progress in this area. To bridge this gap, we present RouterBench, a novel evaluation framework designed to systematically assess the efficacy of LLM routing systems, along with a comprehensive dataset comprising over 405k inference outcomes from representative LLMs to support the development of routing strategies. We further propose a theoretical framework for LLM routing, and deliver a comparative analysis of various routing approaches through RouterBench, highlighting their potentials and limitations within our evaluation framework. This work not only formalizes and advances the development of LLM routing systems but also sets a standard for their assessment, paving the way for more accessible and economically viable LLM deployments. The code and data are available at https://github.com/withmartian/routerbench.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12031
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RouterBench: A Benchmark for Multi-LLM Routing System
Hu, Qitian Jason
Bieker, Jacob
Li, Xiuyu
Jiang, Nan
Keigwin, Benjamin
Ranganath, Gaurav
Keutzer, Kurt
Upadhyay, Shriyash Kaustubh
Machine Learning
Artificial Intelligence
As the range of applications for Large Language Models (LLMs) continues to grow, the demand for effective serving solutions becomes increasingly critical. Despite the versatility of LLMs, no single model can optimally address all tasks and applications, particularly when balancing performance with cost. This limitation has led to the development of LLM routing systems, which combine the strengths of various models to overcome the constraints of individual LLMs. Yet, the absence of a standardized benchmark for evaluating the performance of LLM routers hinders progress in this area. To bridge this gap, we present RouterBench, a novel evaluation framework designed to systematically assess the efficacy of LLM routing systems, along with a comprehensive dataset comprising over 405k inference outcomes from representative LLMs to support the development of routing strategies. We further propose a theoretical framework for LLM routing, and deliver a comparative analysis of various routing approaches through RouterBench, highlighting their potentials and limitations within our evaluation framework. This work not only formalizes and advances the development of LLM routing systems but also sets a standard for their assessment, paving the way for more accessible and economically viable LLM deployments. The code and data are available at https://github.com/withmartian/routerbench.
title RouterBench: A Benchmark for Multi-LLM Routing System
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2403.12031