vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xunzhuo, Chen, Huamin, Lu, Samzong, Ovadia, Yossi, Wen, Guohong, Wu, Hao, Tan, Zhengda, Zhang, Jintao, Zedan, Senan, Kerido, Yehudit, Weiss, Liav, Zhang, Haichen, Yu, Bishen, Balum, Asaad, Limoy, Noa, Samara, Abdallah, Fan, Baofa, Salisbury, Brent, Cook, Ryan, Wang, Zhijie, Pan, Qiping, Khan, Rehan, Goswami, Avishek, Zhang, Houston H., Wang, Shuyi, Tang, Ziang, Han, Fang, Hassan, Zohaib, Zheng, Jianqiao, Changrani, Avinash |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
When to Reason: Semantic Router for vLLM
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Fast and Faithful: Real-Time Verification for Long-Document Retrieval-Augmented Generation Systems
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Adaptive Vision-Language Model Routing for Computer Use Agents
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Knowledge Access Beats Model Size: Memory Augmented Routing for Persistent AI Agents
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Conflict-Free Policy Languages for Probabilistic ML Predicates: A Framework and Case Study with the Semantic Router DSL
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)
Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models
von: Tsfaty, Noam, et al.
Veröffentlicht: (2025)
von: Tsfaty, Noam, et al.
Veröffentlicht: (2025)
From Inference Routing to Agent Orchestration: Declarative Policy Compilation with Cross-Layer Verification
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2026)
von: Ko, Ching-Yun, et al.
Veröffentlicht: (2026)
Counting words without strictly increasing subwords of fixed length
von: Sekhon, Senan
Veröffentlicht: (2025)
von: Sekhon, Senan
Veröffentlicht: (2025)
A Necessary and Sufficient Condition for Uniqueness of Euclidean Division
von: Sekhon, Senan
Veröffentlicht: (2026)
von: Sekhon, Senan
Veröffentlicht: (2026)
LA EXCEPCIÓN EN EL DERECHO. DISCUSIÓN DEL ESTADO DE EXCEPCIÓN EN LA TEORÍA JURÍDICO POLÍTICA
von: Marcela Chahuán Zedán
Veröffentlicht: (2013)
von: Marcela Chahuán Zedán
Veröffentlicht: (2013)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Central Limit Theorem for Mutation Systems
von: Koram, Liav, et al.
Veröffentlicht: (2026)
von: Koram, Liav, et al.
Veröffentlicht: (2026)
Context-Aware Knowledge Distillation with Adaptive Weighting for Image Classification
von: Li, Zhengda
Veröffentlicht: (2025)
von: Li, Zhengda
Veröffentlicht: (2025)
Endogenous knowledge and ethnobotanical importance of Treculia africana Decne. ssp. var. africana in southern Benin (West Africa)
von: Pascaline Sènan, Davoudou
Veröffentlicht: (2024)
von: Pascaline Sènan, Davoudou
Veröffentlicht: (2024)
Causal Mediation in Natural Experiments
von: Hogan-Hennessy, Senan
Veröffentlicht: (2025)
von: Hogan-Hennessy, Senan
Veröffentlicht: (2025)
Consumo y cambio social en España: evolución en el equipamiento doméstico (1983-2005)
von: Gaspar Brändle Señán
Veröffentlicht: (2007)
von: Gaspar Brändle Señán
Veröffentlicht: (2007)
EL CONSUMO EN TIEMPOS DE CRISIS: UNA APROXIMACIÓN SOCIOLÓGICA A LA DISTRIBUCIÓN DEL GASTO EN ESPAÑA
von: Gaspar Brändle Señán
Veröffentlicht: (2010)
von: Gaspar Brändle Señán
Veröffentlicht: (2010)
The Benard-Conway invariant of two-component links
von: Liu, Zedan, et al.
Veröffentlicht: (2024)
von: Liu, Zedan, et al.
Veröffentlicht: (2024)
Exploring the Potential of Twitter as a Research Tool
von: Ovadia, Steven
Veröffentlicht: (2009)
von: Ovadia, Steven
Veröffentlicht: (2009)
Quora.com: Another Place for Users to Ask Questions
von: Ovadia, Steven
Veröffentlicht: (2011)
von: Ovadia, Steven
Veröffentlicht: (2011)
The Viability of Google Wave as an Online Collaboration Tool
von: Ovadia, Steven
Veröffentlicht: (2010)
von: Ovadia, Steven
Veröffentlicht: (2010)
The Role of Big Data in the Social Sciences
von: Ovadia, Steven
Veröffentlicht: (2013)
von: Ovadia, Steven
Veröffentlicht: (2013)
How Does Tenure Status Impact Library Usage: A Study of LaGuardia Community College
von: Ovadia, Steven
Veröffentlicht: (2009)
von: Ovadia, Steven
Veröffentlicht: (2009)
Working without a Crystal Ball: Predicting Web Trends for Web Services Librarians
von: Ovadia, Steven
Veröffentlicht: (2008)
von: Ovadia, Steven
Veröffentlicht: (2008)
Chemical Composition Regulation for Tuning the Luminescent Properties of Organic–Inorganic Metal Halides
von: Zhizhuan Zhang, et al.
Veröffentlicht: (2025)
von: Zhizhuan Zhang, et al.
Veröffentlicht: (2025)
RouterArena: An Open Platform for Comprehensive Comparison of LLM Routers
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
von: Lu, Yifan, et al.
Veröffentlicht: (2025)
Mixture of Routers
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Jia-Chen, et al.
Veröffentlicht: (2025)
TCAndon-Router: Adaptive Reasoning Router for Multi-Agent Collaboration
von: Zhao, Jiuzhou, et al.
Veröffentlicht: (2026)
von: Zhao, Jiuzhou, et al.
Veröffentlicht: (2026)
Router Upcycling: Leveraging Mixture-of-Routers in Mixture-of-Experts Upcycling
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
von: Ran, Junfeng, et al.
Veröffentlicht: (2025)
Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
von: He, Eric, et al.
Veröffentlicht: (2025)
von: He, Eric, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026) -
When to Reason: Semantic Router for vLLM
von: Wang, Chen, et al.
Veröffentlicht: (2025) -
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
von: Chen, Huamin, et al.
Veröffentlicht: (2026) -
Fast and Faithful: Real-Time Verification for Long-Document Retrieval-Augmented Generation Systems
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026) -
Adaptive Vision-Language Model Routing for Computer Use Agents
von: Liu, Xunzhuo, et al.
Veröffentlicht: (2026)