BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ding, Dujian, Mallick, Ankur, Zhang, Shaokun, Wang, Chi, Madrigal, Daniel, Garcia, Mirian Del Carmen Hipolito, Xia, Menglin, Lakshmanan, Laks V. S., Wu, Qingyun, Rühle, Victor
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911027428327424
author Ding, Dujian
Mallick, Ankur
Zhang, Shaokun
Wang, Chi
Madrigal, Daniel
Garcia, Mirian Del Carmen Hipolito
Xia, Menglin
Lakshmanan, Laks V. S.
Wu, Qingyun
Rühle, Victor
author_facet Ding, Dujian
Mallick, Ankur
Zhang, Shaokun
Wang, Chi
Madrigal, Daniel
Garcia, Mirian Del Carmen Hipolito
Xia, Menglin
Lakshmanan, Laks V. S.
Wu, Qingyun
Rühle, Victor
contents Large language models (LLMs) are powerful tools but are often expensive to deploy at scale. LLM query routing mitigates this by dynamically assigning queries to models of varying cost and quality to obtain a desired trade-off. Prior query routing approaches generate only one response from the selected model and a single response from a small (inexpensive) model was often not good enough to beat a response from a large (expensive) model due to which they end up overusing the large model and missing out on potential cost savings. However, it is well known that for small models, generating multiple responses and selecting the best can enhance quality while remaining cheaper than a single large-model response. We leverage this idea to propose BEST-Route, a novel routing framework that chooses a model and the number of responses to sample from it based on query difficulty and the quality thresholds. Experiments on real-world datasets demonstrate that our method reduces costs by up to 60% with less than 1% performance drop.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22716
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
Ding, Dujian
Mallick, Ankur
Zhang, Shaokun
Wang, Chi
Madrigal, Daniel
Garcia, Mirian Del Carmen Hipolito
Xia, Menglin
Lakshmanan, Laks V. S.
Wu, Qingyun
Rühle, Victor
Machine Learning
Artificial Intelligence
Computation and Language
Databases
Large language models (LLMs) are powerful tools but are often expensive to deploy at scale. LLM query routing mitigates this by dynamically assigning queries to models of varying cost and quality to obtain a desired trade-off. Prior query routing approaches generate only one response from the selected model and a single response from a small (inexpensive) model was often not good enough to beat a response from a large (expensive) model due to which they end up overusing the large model and missing out on potential cost savings. However, it is well known that for small models, generating multiple responses and selecting the best can enhance quality while remaining cheaper than a single large-model response. We leverage this idea to propose BEST-Route, a novel routing framework that chooses a model and the number of responses to sample from it based on query difficulty and the quality thresholds. Experiments on real-world datasets demonstrate that our method reduces costs by up to 60% with less than 1% performance drop.
title BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
topic Machine Learning
Artificial Intelligence
Computation and Language
Databases
url https://arxiv.org/abs/2506.22716