Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Kexin, Xiang, Dawei, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
by: Chu, Kexin
Published: (2026)
by: Chu, Kexin
Published: (2026)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
by: Chu, Kexin, et al.
Published: (2026)
by: Chu, Kexin, et al.
Published: (2026)
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
by: Bao, Rui, et al.
Published: (2025)
by: Bao, Rui, et al.
Published: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
Stateful Inference for Low-Latency Multi-Agent Tool Calling
by: Norgren, Victor
Published: (2026)
by: Norgren, Victor
Published: (2026)
SkillRouter: Skill Routing for LLM Agents at Scale
by: Zheng, YanZhao, et al.
Published: (2026)
by: Zheng, YanZhao, et al.
Published: (2026)
MasRouter: Learning to Route LLMs for Multi-Agent Systems
by: Yue, Yanwei, et al.
Published: (2025)
by: Yue, Yanwei, et al.
Published: (2025)
AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive Reasoning
by: Wu, Shirley, et al.
Published: (2024)
by: Wu, Shirley, et al.
Published: (2024)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
PLAY2PROMPT: Zero-shot Tool Instruction Optimization for LLM Agents via Tool Play
by: Fang, Wei, et al.
Published: (2025)
by: Fang, Wei, et al.
Published: (2025)
Memory-Induced Tool-Drift in LLM Agents
by: Dabas, Mahavir, et al.
Published: (2026)
by: Dabas, Mahavir, et al.
Published: (2026)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
by: Acikgoz, Emre Can, et al.
Published: (2026)
by: Acikgoz, Emre Can, et al.
Published: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
LLM Agents Making Agent Tools
by: Wölflein, Georg, et al.
Published: (2025)
by: Wölflein, Georg, et al.
Published: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
by: Xiong, Weimin, et al.
Published: (2025)
by: Xiong, Weimin, et al.
Published: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
by: Zhang, Wenbo, et al.
Published: (2026)
by: Zhang, Wenbo, et al.
Published: (2026)
LLM Routing with Dueling Feedback
by: Chiang, Chao-Kai, et al.
Published: (2025)
by: Chiang, Chao-Kai, et al.
Published: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning
by: Xu, Zhongling, et al.
Published: (2026)
by: Xu, Zhongling, et al.
Published: (2026)
TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents
by: Kumar, Abhishek Vijaya, et al.
Published: (2026)
by: Kumar, Abhishek Vijaya, et al.
Published: (2026)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
by: Winston, Caleb, et al.
Published: (2026)
by: Winston, Caleb, et al.
Published: (2026)
When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
RouteLLM: Learning to Route LLMs with Preference Data
by: Ong, Isaac, et al.
Published: (2024)
by: Ong, Isaac, et al.
Published: (2024)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
by: Du, Yin, et al.
Published: (2026)
by: Du, Yin, et al.
Published: (2026)
Open World Learning Graph Convolution for Latency Estimation in Routing Networks
by: Jin, Yifei, et al.
Published: (2022)
by: Jin, Yifei, et al.
Published: (2022)
MetaToolAgent: Towards Generalizable Tool Usage in LLMs through Meta-Learning
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
by: Thaman, Kunvar
Published: (2026)
by: Thaman, Kunvar
Published: (2026)
Adaptive LLM Routing under Budget Constraints
by: Panda, Pranoy, et al.
Published: (2025)
by: Panda, Pranoy, et al.
Published: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
ToolACE: Winning the Points of LLM Function Calling
by: Liu, Weiwen, et al.
Published: (2024)
by: Liu, Weiwen, et al.
Published: (2024)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
by: Lai, Kunfeng, et al.
Published: (2025)
by: Lai, Kunfeng, et al.
Published: (2025)
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
by: Huang, Hanxian, et al.
Published: (2026)
by: Huang, Hanxian, et al.
Published: (2026)
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
by: Xiang, Zhen, et al.
Published: (2024)
by: Xiang, Zhen, et al.
Published: (2024)
GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
by: Regmi, Sajal, et al.
Published: (2024)
by: Regmi, Sajal, et al.
Published: (2024)
Route Sparse Autoencoder to Interpret Large Language Models
by: Shi, Wei, et al.
Published: (2025)
by: Shi, Wei, et al.
Published: (2025)
Similar Items
-
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
by: Chu, Kexin
Published: (2026) -
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
by: Chu, Kexin, et al.
Published: (2026) -
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
by: Bao, Rui, et al.
Published: (2025) -
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
by: Chu, Kexin, et al.
Published: (2025) -
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)