TensorOpera Router: A Multi-Model Router for Efficient LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Stripelis, Dimitris, Hu, Zijian, Zhang, Jipeng, Xu, Zhaozhuo, Shah, Alay Dilipbhai, Jin, Han, Yao, Yuhang, Avestimehr, Salman, He, Chaoyang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TorchOpera: A Compound AI System for LLM Safety
by: Han, Shanshan, et al.
Published: (2024)
by: Han, Shanshan, et al.
Published: (2024)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
Fox-1: Open Small Language Model for Cloud and Edge
by: Hu, Zijian, et al.
Published: (2024)
by: Hu, Zijian, et al.
Published: (2024)
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
by: Ran, Yide, et al.
Published: (2024)
by: Ran, Yide, et al.
Published: (2024)
One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
by: Shakirov, Georgiy, et al.
Published: (2025)
by: Shakirov, Georgiy, et al.
Published: (2025)
The Federation Strikes Back: A Survey of Federated Learning Privacy Attacks, Defenses, Applications, and Policy Landscape
by: Zhao, Joshua C., et al.
Published: (2024)
by: Zhao, Joshua C., et al.
Published: (2024)
Chain-Oriented Objective Logic with Neural Network Feedback Control and Cascade Filtering for Dynamic Multi-DSL Regulation
by: Han, Jipeng
Published: (2024)
by: Han, Jipeng
Published: (2024)
Regime-Conditional Retrieval: Theory and a Transferable Router for Two-Hop QA
by: Bacellar, Andre
Published: (2026)
by: Bacellar, Andre
Published: (2026)
Toward Super Agent System with Hybrid AI Routers
by: Yao, Yuhang, et al.
Published: (2025)
by: Yao, Yuhang, et al.
Published: (2025)
EM-Aware Physical Synthesis: Neural Inductor Modeling and Intelligent Placement & Routing for RF Circuits
by: Huang, Yilun, et al.
Published: (2026)
by: Huang, Yilun, et al.
Published: (2026)
ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
by: Zhu, Xiongwei, et al.
Published: (2026)
by: Zhu, Xiongwei, et al.
Published: (2026)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
by: Ashley, Dylan R., et al.
Published: (2026)
by: Ashley, Dylan R., et al.
Published: (2026)
Tensor Logic: The Language of AI
by: Domingos, Pedro
Published: (2025)
by: Domingos, Pedro
Published: (2025)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
A Resource-Efficient Hybrid CNN-LSTM network for image-based bean leaf disease classification
by: Rhee, Hye Jin, et al.
Published: (2026)
by: Rhee, Hye Jin, et al.
Published: (2026)
FlexQuant: A Flexible and Efficient Dynamic Precision Switching Framework for LLM Quantization
by: Liu, Fangxin, et al.
Published: (2025)
by: Liu, Fangxin, et al.
Published: (2025)
Efficient Strategy for Improving Large Language Model (LLM) Capabilities
by: Gutiérrez, Julián Camilo Velandia
Published: (2025)
by: Gutiérrez, Julián Camilo Velandia
Published: (2025)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)
by: Atad, Ido Andrew, et al.
Published: (2026)
Neural Router: Semantic Content Matching for Agentic AI
by: Lovén, Lauri, et al.
Published: (2026)
by: Lovén, Lauri, et al.
Published: (2026)
Intelligence Inertia: Physical Isomorphism and Applications
by: Han, Jipeng
Published: (2026)
by: Han, Jipeng
Published: (2026)
MESS+: Dynamically Learned Inference-Time LLM Routing in Model Zoos with Service Level Guarantees
by: Woisetschläger, Herbert, et al.
Published: (2025)
by: Woisetschläger, Herbert, et al.
Published: (2025)
JotlasNet: Joint Tensor Low-Rank and Attention-based Sparse Unrolling Network for Accelerating Dynamic MRI
by: Zhang, Yinghao, et al.
Published: (2025)
by: Zhang, Yinghao, et al.
Published: (2025)
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
by: Han, Shanshan, et al.
Published: (2025)
by: Han, Shanshan, et al.
Published: (2025)
VERITAS-NLI : Validation and Extraction of Reliable Information Through Automated Scraping and Natural Language Inference
by: Shah, Arjun, et al.
Published: (2024)
by: Shah, Arjun, et al.
Published: (2024)
FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
by: Wang, Zixing, et al.
Published: (2025)
by: Wang, Zixing, et al.
Published: (2025)
Efficient Differentiable Hardware Rasterization for 3D Gaussian Splatting
by: Yuan, Yitian, et al.
Published: (2025)
by: Yuan, Yitian, et al.
Published: (2025)
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
Inducing Semi-Structured Sparsity by Masking for Efficient Model Inference in Convolutional Networks
by: Danhofer, David A.
Published: (2024)
by: Danhofer, David A.
Published: (2024)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
by: Yang, Jisoo, et al.
Published: (2026)
by: Yang, Jisoo, et al.
Published: (2026)
Inverse 3D Microscopy Rendering for Cell Shape Inference with Active Mesh
by: Ichbiah, Sacha, et al.
Published: (2023)
by: Ichbiah, Sacha, et al.
Published: (2023)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
by: Marmoret, Axel, et al.
Published: (2025)
by: Marmoret, Axel, et al.
Published: (2025)
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
by: Lu, Hongyu, et al.
Published: (2026)
by: Lu, Hongyu, et al.
Published: (2026)
Tensor Generalized Approximate Message Passing
by: Li, Yinchuan, et al.
Published: (2025)
by: Li, Yinchuan, et al.
Published: (2025)
Composition of rocks and minerals from the Mid-Atlantic Ridge 5-7°N
by: Pushcharovsky, Yury M, et al.
Published: (2004)
by: Pushcharovsky, Yury M, et al.
Published: (2004)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Chemical composition of basalts and basaltic glasses from the Sierra Leone Fracture Zone region
by: Skolotnev, Sergey G, et al.
Published: (2003)
by: Skolotnev, Sergey G, et al.
Published: (2003)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
by: Du, Bangde, et al.
Published: (2025)
by: Du, Bangde, et al.
Published: (2025)
Cyclic Contrastive Knowledge Transfer for Open-Vocabulary Object Detection
by: Zhang, Chuhan, et al.
Published: (2025)
by: Zhang, Chuhan, et al.
Published: (2025)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
by: Grönquist, Peter, et al.
Published: (2023)
by: Grönquist, Peter, et al.
Published: (2023)
Physical oceanography during L' Atalante cruise Almofront-1
by: Prieur, Louis Marie
Published: (2012)
by: Prieur, Louis Marie
Published: (2012)
Similar Items
-
TorchOpera: A Compound AI System for LLM Safety
by: Han, Shanshan, et al.
Published: (2024) -
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024) -
Fox-1: Open Small Language Model for Cloud and Edge
by: Hu, Zijian, et al.
Published: (2024) -
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
by: Ran, Yide, et al.
Published: (2024) -
One Router to Route Them All: Homogeneous Expert Routing for Heterogeneous Graph Transformers
by: Shakirov, Georgiy, et al.
Published: (2025)