Towards Cost-effective LLMs Routing with Batch Prompting
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Haotian, Zhao, Kangfei, Xie, Jiadong |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Beyond Linear LLM Invocation: An Efficient and Effective Semantic Filter Paradigm
par: Hou, Nan, et autres
Publié: (2026)
par: Hou, Nan, et autres
Publié: (2026)
CardOOD: Robust Query-driven Cardinality Estimation under Out-of-Distribution
par: Li, Rui, et autres
Publié: (2024)
par: Li, Rui, et autres
Publié: (2024)
Sema: A High-performance System for LLM-based Semantic Query Processing
par: Qi, Kangkang, et autres
Publié: (2026)
par: Qi, Kangkang, et autres
Publié: (2026)
Can Large Language Models Be Query Optimizer for Relational Databases?
par: Tan, Jie, et autres
Publié: (2025)
par: Tan, Jie, et autres
Publié: (2025)
Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]
par: Liu, Yingfan, et autres
Publié: (2026)
par: Liu, Yingfan, et autres
Publié: (2026)
Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional Data
par: Liu, Yingfan, et autres
Publié: (2025)
par: Liu, Yingfan, et autres
Publié: (2025)
Towards Effective Orchestration of AI x DB Workloads
par: Xing, Naili, et autres
Publié: (2026)
par: Xing, Naili, et autres
Publié: (2026)
Revisiting the Index Construction of Proximity Graph-Based Approximate Nearest Neighbor Search
par: Yang, Shuo, et autres
Publié: (2024)
par: Yang, Shuo, et autres
Publié: (2024)
Towards Optimizing SQL Generation via LLM Routing
par: Malekpour, Mohammadhossein, et autres
Publié: (2024)
par: Malekpour, Mohammadhossein, et autres
Publié: (2024)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
par: Kassem, Aly M., et autres
Publié: (2025)
par: Kassem, Aly M., et autres
Publié: (2025)
Fast Tuning the Index Construction Parameters of Proximity Graphs in Vector Databases
par: Zhou, Wenyang, et autres
Publié: (2026)
par: Zhou, Wenyang, et autres
Publié: (2026)
BRkNN-light: Batch Processing of Reverse k-Nearest Neighbor Queries for Moving Objects on Road Networks
par: Song, Anbang, et autres
Publié: (2025)
par: Song, Anbang, et autres
Publié: (2025)
Towards Scalable and Practical Batch-Dynamic Connectivity
par: De Man, Quinten, et autres
Publié: (2024)
par: De Man, Quinten, et autres
Publié: (2024)
EllieSQL: Cost-Efficient Text-to-SQL with Complexity-Aware Routing
par: Zhu, Yizhang, et autres
Publié: (2025)
par: Zhu, Yizhang, et autres
Publié: (2025)
Batch Hop-Constrained s-t Simple Path Query Processing in Large Graphs
par: Yuan, Long, et autres
Publié: (2023)
par: Yuan, Long, et autres
Publié: (2023)
ScaleGANN: Accelerate Large-Scale ANN Indexing by Cost-effective Cloud GPUs
par: Lu, Lan, et autres
Publié: (2026)
par: Lu, Lan, et autres
Publié: (2026)
BQSched: A Non-intrusive Scheduler for Batch Concurrent Queries via Reinforcement Learning
par: Xu, Chenhao, et autres
Publié: (2025)
par: Xu, Chenhao, et autres
Publié: (2025)
TrajRoute: Rethinking Routing with a Simple Trajectory-Based Approach -- Forget the Maps and Traffic!
par: Siampou, Maria Despoina, et autres
Publié: (2024)
par: Siampou, Maria Despoina, et autres
Publié: (2024)
GeoLayer: Towards Low-Latency and Cost-Efficient Geo-Distributed Graph Stores with Layered Graph
par: Yao, Feng, et autres
Publié: (2025)
par: Yao, Feng, et autres
Publié: (2025)
Cost-Efficient RAG for Entity Matching with LLMs: A Blocking-based Exploration
par: Ma, Chuangtao, et autres
Publié: (2026)
par: Ma, Chuangtao, et autres
Publié: (2026)
Influence Minimization via Blocking Strategies
par: Xie, Jiadong, et autres
Publié: (2023)
par: Xie, Jiadong, et autres
Publié: (2023)
Structured Prompt Language: Declarative Context Management for LLMs
par: Gong, Wen G.
Publié: (2026)
par: Gong, Wen G.
Publié: (2026)
Bounding the Fragmentation of B-Trees Subject to Batched Insertions
par: Bender, Michael A., et autres
Publié: (2026)
par: Bender, Michael A., et autres
Publié: (2026)
Flexible Keyword-Aware Top-$k$ Route Search
par: Yu, Ziqiang, et autres
Publié: (2025)
par: Yu, Ziqiang, et autres
Publié: (2025)
The Cost of Representation by Subset Repairs
par: Liu, Yuxi, et autres
Publié: (2024)
par: Liu, Yuxi, et autres
Publié: (2024)
Towards Temporal Knowledge Graph Alignment in the Wild
par: Zhao, Runhao, et autres
Publié: (2025)
par: Zhao, Runhao, et autres
Publié: (2025)
OmniRouter: Budget and Performance Controllable Multi-LLM Routing
par: Mei, Kai, et autres
Publié: (2025)
par: Mei, Kai, et autres
Publié: (2025)
Batch Query Processing and Optimization for Agentic Workflows
par: Shen, Junyi, et autres
Publié: (2025)
par: Shen, Junyi, et autres
Publié: (2025)
Time Sensitive Multiple POIs Route Planning on Bus Networks
par: Liu, Simu, et autres
Publié: (2025)
par: Liu, Simu, et autres
Publié: (2025)
COLE$^+$: Towards Practical Column-based Learned Storage for Blockchain Systems
par: Zhang, Ce, et autres
Publié: (2026)
par: Zhang, Ce, et autres
Publié: (2026)
GRACEFUL: A Learned Cost Estimator For UDFs
par: Wehrstein, Johannes, et autres
Publié: (2025)
par: Wehrstein, Johannes, et autres
Publié: (2025)
PRIME: Efficient Algorithm for Token Graph Routing Problem
par: Xu, Haotian, et autres
Publié: (2026)
par: Xu, Haotian, et autres
Publié: (2026)
From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
par: Feng, Yichao, et autres
Publié: (2025)
par: Feng, Yichao, et autres
Publié: (2025)
Natural Language Interfaces for Databases: What Do Users Think?
par: Ipeirotis, Panos, et autres
Publié: (2025)
par: Ipeirotis, Panos, et autres
Publié: (2025)
HIGGS: HIerarchy-Guided Graph Stream Summarization
par: Zhao, Xuan, et autres
Publié: (2024)
par: Zhao, Xuan, et autres
Publié: (2024)
Cost-based Selection of Provenance Sketches for Data Skipping
par: Liu, Ziyu, et autres
Publié: (2025)
par: Liu, Ziyu, et autres
Publié: (2025)
NeurBench: A Benchmark Suite for Learned Database Components with Drift Modeling
par: Zhao, Zhanhao, et autres
Publié: (2025)
par: Zhao, Zhanhao, et autres
Publié: (2025)
Efficient Cost-Based Rewrite in a Bottom-Up Optimizer
par: Cheng, Qi, et autres
Publié: (2026)
par: Cheng, Qi, et autres
Publié: (2026)
BatchBench: Toward a Workload-Aware Benchmark for Autoscaling Policies in Big Data Batch Processing -- A Proposed Framework
par: Budigi, Venkata Krishna Prasanth, et autres
Publié: (2026)
par: Budigi, Venkata Krishna Prasanth, et autres
Publié: (2026)
Opening The Black-Box: Explaining Learned Cost Models For Databases
par: Heinrich, Roman, et autres
Publié: (2025)
par: Heinrich, Roman, et autres
Publié: (2025)
Documents similaires
-
Beyond Linear LLM Invocation: An Efficient and Effective Semantic Filter Paradigm
par: Hou, Nan, et autres
Publié: (2026) -
CardOOD: Robust Query-driven Cardinality Estimation under Out-of-Distribution
par: Li, Rui, et autres
Publié: (2024) -
Sema: A High-performance System for LLM-based Semantic Query Processing
par: Qi, Kangkang, et autres
Publié: (2026) -
Can Large Language Models Be Query Optimizer for Relational Databases?
par: Tan, Jie, et autres
Publié: (2025) -
Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]
par: Liu, Yingfan, et autres
Publié: (2026)