When to Reason: Semantic Router for vLLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915543254040576 |
|---|---|
| author | Wang, Chen Liu, Xunzhuo Liu, Yuhan Zhu, Yue Mo, Xiangxi Jiang, Junchen Chen, Huamin |
| author_facet | Wang, Chen Liu, Xunzhuo Liu, Yuhan Zhu, Yue Mo, Xiangxi Jiang, Junchen Chen, Huamin |
| contents | Large Language Models (LLMs) demonstrate substantial accuracy gains when augmented with reasoning modes such as chain-of-thought and inference-time scaling. However, reasoning also incurs significant costs in inference latency and token usage, with environmental and financial impacts, which are unnecessary for many simple prompts. We present a semantic router that classifies queries based on their reasoning requirements and selectively applies reasoning only when beneficial. Our approach achieves a 10.2 percentage point improvement in accuracy on the MMLU-Pro benchmark while reducing response latency by 47.1% and token consumption by 48.5% compared to direct inference with vLLM. These results demonstrate that semantic routing offers an effective mechanism for striking a balance between accuracy and efficiency in open-source LLM serving systems |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08731 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | When to Reason: Semantic Router for vLLM Wang, Chen Liu, Xunzhuo Liu, Yuhan Zhu, Yue Mo, Xiangxi Jiang, Junchen Chen, Huamin Emerging Technologies Artificial Intelligence Computation and Language Systems and Control Large Language Models (LLMs) demonstrate substantial accuracy gains when augmented with reasoning modes such as chain-of-thought and inference-time scaling. However, reasoning also incurs significant costs in inference latency and token usage, with environmental and financial impacts, which are unnecessary for many simple prompts. We present a semantic router that classifies queries based on their reasoning requirements and selectively applies reasoning only when beneficial. Our approach achieves a 10.2 percentage point improvement in accuracy on the MMLU-Pro benchmark while reducing response latency by 47.1% and token consumption by 48.5% compared to direct inference with vLLM. These results demonstrate that semantic routing offers an effective mechanism for striking a balance between accuracy and efficiency in open-source LLM serving systems |
| title | When to Reason: Semantic Router for vLLM |
| topic | Emerging Technologies Artificial Intelligence Computation and Language Systems and Control |
| url | https://arxiv.org/abs/2510.08731 |