When to Reason: Semantic Router for vLLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chen, Liu, Xunzhuo, Liu, Yuhan, Zhu, Yue, Mo, Xiangxi, Jiang, Junchen, Chen, Huamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915543254040576
author Wang, Chen
Liu, Xunzhuo
Liu, Yuhan
Zhu, Yue
Mo, Xiangxi
Jiang, Junchen
Chen, Huamin
author_facet Wang, Chen
Liu, Xunzhuo
Liu, Yuhan
Zhu, Yue
Mo, Xiangxi
Jiang, Junchen
Chen, Huamin
contents Large Language Models (LLMs) demonstrate substantial accuracy gains when augmented with reasoning modes such as chain-of-thought and inference-time scaling. However, reasoning also incurs significant costs in inference latency and token usage, with environmental and financial impacts, which are unnecessary for many simple prompts. We present a semantic router that classifies queries based on their reasoning requirements and selectively applies reasoning only when beneficial. Our approach achieves a 10.2 percentage point improvement in accuracy on the MMLU-Pro benchmark while reducing response latency by 47.1% and token consumption by 48.5% compared to direct inference with vLLM. These results demonstrate that semantic routing offers an effective mechanism for striking a balance between accuracy and efficiency in open-source LLM serving systems
format Preprint
id arxiv_https___arxiv_org_abs_2510_08731
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When to Reason: Semantic Router for vLLM
Wang, Chen
Liu, Xunzhuo
Liu, Yuhan
Zhu, Yue
Mo, Xiangxi
Jiang, Junchen
Chen, Huamin
Emerging Technologies
Artificial Intelligence
Computation and Language
Systems and Control
Large Language Models (LLMs) demonstrate substantial accuracy gains when augmented with reasoning modes such as chain-of-thought and inference-time scaling. However, reasoning also incurs significant costs in inference latency and token usage, with environmental and financial impacts, which are unnecessary for many simple prompts. We present a semantic router that classifies queries based on their reasoning requirements and selectively applies reasoning only when beneficial. Our approach achieves a 10.2 percentage point improvement in accuracy on the MMLU-Pro benchmark while reducing response latency by 47.1% and token consumption by 48.5% compared to direct inference with vLLM. These results demonstrate that semantic routing offers an effective mechanism for striking a balance between accuracy and efficiency in open-source LLM serving systems
title When to Reason: Semantic Router for vLLM
topic Emerging Technologies
Artificial Intelligence
Computation and Language
Systems and Control
url https://arxiv.org/abs/2510.08731