RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hao, Sai, Zeng, Hao, Wei, Hongxin, Jing, Bingyi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918376074379264
author Hao, Sai
Zeng, Hao
Wei, Hongxin
Jing, Bingyi
author_facet Hao, Sai
Zeng, Hao
Wei, Hongxin
Jing, Bingyi
contents Efficiently routing queries to the optimal large language model (LLM) is crucial for optimizing the cost-performance trade-off in multi-model systems. However, most existing routers rely on single-model selection, making them susceptible to misrouting. In this work, we formulate LLM routing as the $α$-VOR problem to minimize expected set size while controlling the misrouting risk, and propose a novel method -- RACER, extending base routers to output model sets that can be subsequently aggregated for improved output. In particular, RACER constructs nested model sets via augmented scoring and utilizes finite-sample concentration bounds to calibrate a threshold that allows for both variable set sizes and abstention. We theoretically prove that RACER achieves rigorous distribution-free risk control on unseen test data in a post-hoc and model-agnostic manner. Extensive experiments verify our theoretical guarantees and demonstrate that RACER consistently enhances downstream accuracy across a wide range of benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06616
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
Hao, Sai
Zeng, Hao
Wei, Hongxin
Jing, Bingyi
Machine Learning
Artificial Intelligence
Statistics Theory
Efficiently routing queries to the optimal large language model (LLM) is crucial for optimizing the cost-performance trade-off in multi-model systems. However, most existing routers rely on single-model selection, making them susceptible to misrouting. In this work, we formulate LLM routing as the $α$-VOR problem to minimize expected set size while controlling the misrouting risk, and propose a novel method -- RACER, extending base routers to output model sets that can be subsequently aggregated for improved output. In particular, RACER constructs nested model sets via augmented scoring and utilizes finite-sample concentration bounds to calibrate a threshold that allows for both variable set sizes and abstention. We theoretically prove that RACER achieves rigorous distribution-free risk control on unseen test data in a post-hoc and model-agnostic manner. Extensive experiments verify our theoretical guarantees and demonstrate that RACER consistently enhances downstream accuracy across a wide range of benchmarks.
title RACER: Risk-Aware Calibrated Efficient Routing for Large Language Models
topic Machine Learning
Artificial Intelligence
Statistics Theory
url https://arxiv.org/abs/2603.06616