Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ding, Dujian, Mallick, Ankur, Wang, Chi, Sim, Robert, Mukherjee, Subhabrata, Ruhle, Victor, Lakshmanan, Laks V. S., Awadallah, Ahmed Hassan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OCCAM: Towards Cost-Efficient and Accuracy-Aware Classification Inference
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
von: Ding, Dujian, et al.
Veröffentlicht: (2025)
von: Ding, Dujian, et al.
Veröffentlicht: (2025)
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries
von: Huang, Keke, et al.
Veröffentlicht: (2025)
von: Huang, Keke, et al.
Veröffentlicht: (2025)
LLM Performance Predictors are good initializers for Architecture Search
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023)
EcoAct: Economic Agent Determines When to Register What Action
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)
On Efficient Approximate Aggregate Nearest Neighbor Queries over Learned Representations
von: Wang, Carrie, et al.
Veröffentlicht: (2025)
von: Wang, Carrie, et al.
Veröffentlicht: (2025)
ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
von: Patel, Shivam, et al.
Veröffentlicht: (2025)
von: Patel, Shivam, et al.
Veröffentlicht: (2025)
KRAFT: A Knowledge Graph-Based Framework for Automated Map Conflation
von: Hashemi, Farnoosh, et al.
Veröffentlicht: (2025)
von: Hashemi, Farnoosh, et al.
Veröffentlicht: (2025)
Topology-Aware LLM-Driven Social Simulation: A Unified Framework for Efficient and Realistic Agent Dynamics
von: Xu, Yuwei, et al.
Veröffentlicht: (2026)
von: Xu, Yuwei, et al.
Veröffentlicht: (2026)
Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
von: Hashemi, Helia, et al.
Veröffentlicht: (2025)
von: Hashemi, Helia, et al.
Veröffentlicht: (2025)
Exploring How LLMs Capture and Represent Domain-Specific Knowledge
von: Garcia, Mirian Hipolito, et al.
Veröffentlicht: (2025)
von: Garcia, Mirian Hipolito, et al.
Veröffentlicht: (2025)
DetoxLLM: A Framework for Detoxification with Explanations
von: Khondaker, Md Tawkat Islam, et al.
Veröffentlicht: (2024)
von: Khondaker, Md Tawkat Islam, et al.
Veröffentlicht: (2024)
Sweeping Heterogeneity with Smart MoPs: Mixture of Prompts for LLM Task Adaptation
von: Dun, Chen, et al.
Veröffentlicht: (2023)
von: Dun, Chen, et al.
Veröffentlicht: (2023)
Autoregressive + Chain of Thought = Recurrent: Recurrence's Role in Language Models' Computability and a Revisit of Recurrent Transformer
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
A Multi-Agent Approach for Claim Verification from Tabular Data Documents
von: Saha, Rudra Ranajee, et al.
Veröffentlicht: (2026)
von: Saha, Rudra Ranajee, et al.
Veröffentlicht: (2026)
A Community-Based Approach for Stance Distribution and Argument Organization
von: Saha, Rudra Ranajee, et al.
Veröffentlicht: (2026)
von: Saha, Rudra Ranajee, et al.
Veröffentlicht: (2026)
Supervised Chain of Thought
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
von: Zhang, Xiang, et al.
Veröffentlicht: (2024)
BEACON: A Benchmark for Efficient and Accurate Counting of Subgraphs
von: Najafi, Mohammad Matin, et al.
Veröffentlicht: (2025)
von: Najafi, Mohammad Matin, et al.
Veröffentlicht: (2025)
Budget-Aware Agentic Routing via Boundary-Guided Training
von: Zhang, Caiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Caiqi, et al.
Veröffentlicht: (2026)
A Survey of Densest Subgraph Discovery on Large Graphs
von: Luo, Wensheng, et al.
Veröffentlicht: (2023)
von: Luo, Wensheng, et al.
Veröffentlicht: (2023)
Hyperparametric Robust and Dynamic Influence Maximization
von: Saha, Arkaprava, et al.
Veröffentlicht: (2024)
von: Saha, Arkaprava, et al.
Veröffentlicht: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
SynConfRoute: Syntax-Aware Routing for Efficient Code Completion with Small CodeLLMs
von: Thangarajah, Kishanthan, et al.
Veröffentlicht: (2026)
von: Thangarajah, Kishanthan, et al.
Veröffentlicht: (2026)
Efficient and Effective Algorithms for A Family of Influence Maximization Problems with A Matroid Constraint
von: Huang, Yiqian, et al.
Veröffentlicht: (2024)
von: Huang, Yiqian, et al.
Veröffentlicht: (2024)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Finding Locally Densest Subgraphs: Convex Programming with Edge and Triangle Density
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
TRACE: Intra-visit Clinical Event Nowcasting via Effective Patient Trajectory Encoding
von: Liang, Yuyang, et al.
Veröffentlicht: (2025)
von: Liang, Yuyang, et al.
Veröffentlicht: (2025)
A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and Effectiveness
von: Liao, Ningyi, et al.
Veröffentlicht: (2024)
von: Liao, Ningyi, et al.
Veröffentlicht: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Implementation of Energy‐Aware Optimal Routing for Improving Traffic Capacity in Ad Hoc Wireless Network Using Hybrid Heuristic Algorithm
von: Nandakumar Seekarajapuram Dinakaran, et al.
Veröffentlicht: (2025)
von: Nandakumar Seekarajapuram Dinakaran, et al.
Veröffentlicht: (2025)
Fast Maximum Common Subgraph Search: A Redundancy-Reduced Backtracking Approach
von: Yu, Kaiqiang, et al.
Veröffentlicht: (2025)
von: Yu, Kaiqiang, et al.
Veröffentlicht: (2025)
Cost-Aware Routing for Efficient Text-To-Image Generation
von: Li, Qinchan, et al.
Veröffentlicht: (2025)
von: Li, Qinchan, et al.
Veröffentlicht: (2025)
Adaptive and Robust Cost-Aware Proof of Quality for Decentralized LLM Inference Networks
von: Tian, Arther, et al.
Veröffentlicht: (2026)
von: Tian, Arther, et al.
Veröffentlicht: (2026)
Towards better Human-Agent Alignment: Assessing Task Utility in LLM-Powered Applications
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
von: Arabzadeh, Negar, et al.
Veröffentlicht: (2024)
Predicting Cascading Failures with a Hyperparametric Diffusion Model
von: Xiang, Bin, et al.
Veröffentlicht: (2024)
von: Xiang, Bin, et al.
Veröffentlicht: (2024)
A Multi‐Objective Energy‐Aware Routing in Ad Hoc Networks Using Improved Heuristic Optimization Based on Energy Prediction via Dense Conv‐LSTM With Coordinate Attention
von: S. D. Nandakumar, et al.
Veröffentlicht: (2026)
von: S. D. Nandakumar, et al.
Veröffentlicht: (2026)
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
von: Okamoto, Mika, et al.
Veröffentlicht: (2026)
von: Okamoto, Mika, et al.
Veröffentlicht: (2026)
Markeninszenierung in Japan
von: Rühle, Christiane
Veröffentlicht: (2020)
von: Rühle, Christiane
Veröffentlicht: (2020)
Behavioral and psychological symptoms in dementia is not a unitary concept. A critical review with emphasis on Alzheimer’s disease
von: Jerson Laks
Veröffentlicht: (2008)
von: Jerson Laks
Veröffentlicht: (2008)
Ähnliche Einträge
-
OCCAM: Towards Cost-Efficient and Accuracy-Aware Classification Inference
von: Ding, Dujian, et al.
Veröffentlicht: (2024) -
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
von: Ding, Dujian, et al.
Veröffentlicht: (2025) -
ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification Queries
von: Huang, Keke, et al.
Veröffentlicht: (2025) -
LLM Performance Predictors are good initializers for Architecture Search
von: Jawahar, Ganesh, et al.
Veröffentlicht: (2023) -
EcoAct: Economic Agent Determines When to Register What Action
von: Zhang, Shaokun, et al.
Veröffentlicht: (2024)