Towards Cost-effective LLMs Routing with Batch Prompting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Haotian, Zhao, Kangfei, Xie, Jiadong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Linear LLM Invocation: An Efficient and Effective Semantic Filter Paradigm
von: Hou, Nan, et al.
Veröffentlicht: (2026)
von: Hou, Nan, et al.
Veröffentlicht: (2026)
CardOOD: Robust Query-driven Cardinality Estimation under Out-of-Distribution
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
Sema: A High-performance System for LLM-based Semantic Query Processing
von: Qi, Kangkang, et al.
Veröffentlicht: (2026)
von: Qi, Kangkang, et al.
Veröffentlicht: (2026)
Can Large Language Models Be Query Optimizer for Relational Databases?
von: Tan, Jie, et al.
Veröffentlicht: (2025)
von: Tan, Jie, et al.
Veröffentlicht: (2025)
Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]
von: Liu, Yingfan, et al.
Veröffentlicht: (2026)
von: Liu, Yingfan, et al.
Veröffentlicht: (2026)
Privacy-Preserving Approximate Nearest Neighbor Search on High-Dimensional Data
von: Liu, Yingfan, et al.
Veröffentlicht: (2025)
von: Liu, Yingfan, et al.
Veröffentlicht: (2025)
Towards Effective Orchestration of AI x DB Workloads
von: Xing, Naili, et al.
Veröffentlicht: (2026)
von: Xing, Naili, et al.
Veröffentlicht: (2026)
Revisiting the Index Construction of Proximity Graph-Based Approximate Nearest Neighbor Search
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
Towards Optimizing SQL Generation via LLM Routing
von: Malekpour, Mohammadhossein, et al.
Veröffentlicht: (2024)
von: Malekpour, Mohammadhossein, et al.
Veröffentlicht: (2024)
How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
von: Kassem, Aly M., et al.
Veröffentlicht: (2025)
Fast Tuning the Index Construction Parameters of Proximity Graphs in Vector Databases
von: Zhou, Wenyang, et al.
Veröffentlicht: (2026)
von: Zhou, Wenyang, et al.
Veröffentlicht: (2026)
BRkNN-light: Batch Processing of Reverse k-Nearest Neighbor Queries for Moving Objects on Road Networks
von: Song, Anbang, et al.
Veröffentlicht: (2025)
von: Song, Anbang, et al.
Veröffentlicht: (2025)
Towards Scalable and Practical Batch-Dynamic Connectivity
von: De Man, Quinten, et al.
Veröffentlicht: (2024)
von: De Man, Quinten, et al.
Veröffentlicht: (2024)
EllieSQL: Cost-Efficient Text-to-SQL with Complexity-Aware Routing
von: Zhu, Yizhang, et al.
Veröffentlicht: (2025)
von: Zhu, Yizhang, et al.
Veröffentlicht: (2025)
Batch Hop-Constrained s-t Simple Path Query Processing in Large Graphs
von: Yuan, Long, et al.
Veröffentlicht: (2023)
von: Yuan, Long, et al.
Veröffentlicht: (2023)
ScaleGANN: Accelerate Large-Scale ANN Indexing by Cost-effective Cloud GPUs
von: Lu, Lan, et al.
Veröffentlicht: (2026)
von: Lu, Lan, et al.
Veröffentlicht: (2026)
BQSched: A Non-intrusive Scheduler for Batch Concurrent Queries via Reinforcement Learning
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
von: Xu, Chenhao, et al.
Veröffentlicht: (2025)
TrajRoute: Rethinking Routing with a Simple Trajectory-Based Approach -- Forget the Maps and Traffic!
von: Siampou, Maria Despoina, et al.
Veröffentlicht: (2024)
von: Siampou, Maria Despoina, et al.
Veröffentlicht: (2024)
GeoLayer: Towards Low-Latency and Cost-Efficient Geo-Distributed Graph Stores with Layered Graph
von: Yao, Feng, et al.
Veröffentlicht: (2025)
von: Yao, Feng, et al.
Veröffentlicht: (2025)
Cost-Efficient RAG for Entity Matching with LLMs: A Blocking-based Exploration
von: Ma, Chuangtao, et al.
Veröffentlicht: (2026)
von: Ma, Chuangtao, et al.
Veröffentlicht: (2026)
Influence Minimization via Blocking Strategies
von: Xie, Jiadong, et al.
Veröffentlicht: (2023)
von: Xie, Jiadong, et al.
Veröffentlicht: (2023)
Structured Prompt Language: Declarative Context Management for LLMs
von: Gong, Wen G.
Veröffentlicht: (2026)
von: Gong, Wen G.
Veröffentlicht: (2026)
Bounding the Fragmentation of B-Trees Subject to Batched Insertions
von: Bender, Michael A., et al.
Veröffentlicht: (2026)
von: Bender, Michael A., et al.
Veröffentlicht: (2026)
Flexible Keyword-Aware Top-$k$ Route Search
von: Yu, Ziqiang, et al.
Veröffentlicht: (2025)
von: Yu, Ziqiang, et al.
Veröffentlicht: (2025)
The Cost of Representation by Subset Repairs
von: Liu, Yuxi, et al.
Veröffentlicht: (2024)
von: Liu, Yuxi, et al.
Veröffentlicht: (2024)
Towards Temporal Knowledge Graph Alignment in the Wild
von: Zhao, Runhao, et al.
Veröffentlicht: (2025)
von: Zhao, Runhao, et al.
Veröffentlicht: (2025)
OmniRouter: Budget and Performance Controllable Multi-LLM Routing
von: Mei, Kai, et al.
Veröffentlicht: (2025)
von: Mei, Kai, et al.
Veröffentlicht: (2025)
Batch Query Processing and Optimization for Agentic Workflows
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
Time Sensitive Multiple POIs Route Planning on Bus Networks
von: Liu, Simu, et al.
Veröffentlicht: (2025)
von: Liu, Simu, et al.
Veröffentlicht: (2025)
COLE$^+$: Towards Practical Column-based Learned Storage for Blockchain Systems
von: Zhang, Ce, et al.
Veröffentlicht: (2026)
von: Zhang, Ce, et al.
Veröffentlicht: (2026)
GRACEFUL: A Learned Cost Estimator For UDFs
von: Wehrstein, Johannes, et al.
Veröffentlicht: (2025)
von: Wehrstein, Johannes, et al.
Veröffentlicht: (2025)
PRIME: Efficient Algorithm for Token Graph Routing Problem
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
von: Xu, Haotian, et al.
Veröffentlicht: (2026)
From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
von: Feng, Yichao, et al.
Veröffentlicht: (2025)
von: Feng, Yichao, et al.
Veröffentlicht: (2025)
Natural Language Interfaces for Databases: What Do Users Think?
von: Ipeirotis, Panos, et al.
Veröffentlicht: (2025)
von: Ipeirotis, Panos, et al.
Veröffentlicht: (2025)
HIGGS: HIerarchy-Guided Graph Stream Summarization
von: Zhao, Xuan, et al.
Veröffentlicht: (2024)
von: Zhao, Xuan, et al.
Veröffentlicht: (2024)
Cost-based Selection of Provenance Sketches for Data Skipping
von: Liu, Ziyu, et al.
Veröffentlicht: (2025)
von: Liu, Ziyu, et al.
Veröffentlicht: (2025)
NeurBench: A Benchmark Suite for Learned Database Components with Drift Modeling
von: Zhao, Zhanhao, et al.
Veröffentlicht: (2025)
von: Zhao, Zhanhao, et al.
Veröffentlicht: (2025)
Efficient Cost-Based Rewrite in a Bottom-Up Optimizer
von: Cheng, Qi, et al.
Veröffentlicht: (2026)
von: Cheng, Qi, et al.
Veröffentlicht: (2026)
BatchBench: Toward a Workload-Aware Benchmark for Autoscaling Policies in Big Data Batch Processing -- A Proposed Framework
von: Budigi, Venkata Krishna Prasanth, et al.
Veröffentlicht: (2026)
von: Budigi, Venkata Krishna Prasanth, et al.
Veröffentlicht: (2026)
Opening The Black-Box: Explaining Learned Cost Models For Databases
von: Heinrich, Roman, et al.
Veröffentlicht: (2025)
von: Heinrich, Roman, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond Linear LLM Invocation: An Efficient and Effective Semantic Filter Paradigm
von: Hou, Nan, et al.
Veröffentlicht: (2026) -
CardOOD: Robust Query-driven Cardinality Estimation under Out-of-Distribution
von: Li, Rui, et al.
Veröffentlicht: (2024) -
Sema: A High-performance System for LLM-based Semantic Query Processing
von: Qi, Kangkang, et al.
Veröffentlicht: (2026) -
Can Large Language Models Be Query Optimizer for Relational Databases?
von: Tan, Jie, et al.
Veröffentlicht: (2025) -
Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]
von: Liu, Yingfan, et al.
Veröffentlicht: (2026)