Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chu, Kexin, Xiang, Dawei, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
von: Chu, Kexin
Veröffentlicht: (2026)
von: Chu, Kexin
Veröffentlicht: (2026)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2026)
von: Chu, Kexin, et al.
Veröffentlicht: (2026)
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
von: Bao, Rui, et al.
Veröffentlicht: (2025)
von: Bao, Rui, et al.
Veröffentlicht: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Stateful Inference for Low-Latency Multi-Agent Tool Calling
von: Norgren, Victor
Veröffentlicht: (2026)
von: Norgren, Victor
Veröffentlicht: (2026)
SkillRouter: Skill Routing for LLM Agents at Scale
von: Zheng, YanZhao, et al.
Veröffentlicht: (2026)
von: Zheng, YanZhao, et al.
Veröffentlicht: (2026)
MasRouter: Learning to Route LLMs for Multi-Agent Systems
von: Yue, Yanwei, et al.
Veröffentlicht: (2025)
von: Yue, Yanwei, et al.
Veröffentlicht: (2025)
AvaTaR: Optimizing LLM Agents for Tool Usage via Contrastive Reasoning
von: Wu, Shirley, et al.
Veröffentlicht: (2024)
von: Wu, Shirley, et al.
Veröffentlicht: (2024)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
PLAY2PROMPT: Zero-shot Tool Instruction Optimization for LLM Agents via Tool Play
von: Fang, Wei, et al.
Veröffentlicht: (2025)
von: Fang, Wei, et al.
Veröffentlicht: (2025)
Memory-Induced Tool-Drift in LLM Agents
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026)
von: Dabas, Mahavir, et al.
Veröffentlicht: (2026)
Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2026)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2026)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
LLM Agents Making Agent Tools
von: Wölflein, Georg, et al.
Veröffentlicht: (2025)
von: Wölflein, Georg, et al.
Veröffentlicht: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
von: Wang, Haoxin, et al.
Veröffentlicht: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
von: Ding, Dujian, et al.
Veröffentlicht: (2024)
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
von: Xiong, Weimin, et al.
Veröffentlicht: (2025)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
LLM Routing with Dueling Feedback
von: Chiang, Chao-Kai, et al.
Veröffentlicht: (2025)
von: Chiang, Chao-Kai, et al.
Veröffentlicht: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
von: Kong, Linghao, et al.
Veröffentlicht: (2026)
von: Kong, Linghao, et al.
Veröffentlicht: (2026)
SeqRoute: Global Budget-Aware Sequential LLM Routing via Offline Reinforcement Learning
von: Xu, Zhongling, et al.
Veröffentlicht: (2026)
von: Xu, Zhongling, et al.
Veröffentlicht: (2026)
TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2026)
von: Kumar, Abhishek Vijaya, et al.
Veröffentlicht: (2026)
Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling
von: Winston, Caleb, et al.
Veröffentlicht: (2026)
von: Winston, Caleb, et al.
Veröffentlicht: (2026)
When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
RouteLLM: Learning to Route LLMs with Preference Data
von: Ong, Isaac, et al.
Veröffentlicht: (2024)
von: Ong, Isaac, et al.
Veröffentlicht: (2024)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
von: Gao, Yifei, et al.
Veröffentlicht: (2025)
von: Gao, Yifei, et al.
Veröffentlicht: (2025)
LatencyPrism: Online Non-intrusive Latency Sculpting for SLO-Guaranteed LLM Inference
von: Du, Yin, et al.
Veröffentlicht: (2026)
von: Du, Yin, et al.
Veröffentlicht: (2026)
Open World Learning Graph Convolution for Latency Estimation in Routing Networks
von: Jin, Yifei, et al.
Veröffentlicht: (2022)
von: Jin, Yifei, et al.
Veröffentlicht: (2022)
MetaToolAgent: Towards Generalizable Tool Usage in LLMs through Meta-Learning
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
von: Thaman, Kunvar
Veröffentlicht: (2026)
von: Thaman, Kunvar
Veröffentlicht: (2026)
Adaptive LLM Routing under Budget Constraints
von: Panda, Pranoy, et al.
Veröffentlicht: (2025)
von: Panda, Pranoy, et al.
Veröffentlicht: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
ToolACE: Winning the Points of LLM Function Calling
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
von: Liu, Weiwen, et al.
Veröffentlicht: (2024)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
von: Lai, Kunfeng, et al.
Veröffentlicht: (2025)
von: Lai, Kunfeng, et al.
Veröffentlicht: (2025)
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment
von: Huang, Hanxian, et al.
Veröffentlicht: (2026)
von: Huang, Hanxian, et al.
Veröffentlicht: (2026)
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
von: Xiang, Zhen, et al.
Veröffentlicht: (2024)
GPT Semantic Cache: Reducing LLM Costs and Latency via Semantic Embedding Caching
von: Regmi, Sajal, et al.
Veröffentlicht: (2024)
von: Regmi, Sajal, et al.
Veröffentlicht: (2024)
Route Sparse Autoencoder to Interpret Large Language Models
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
von: Chu, Kexin
Veröffentlicht: (2026) -
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2026) -
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
von: Bao, Rui, et al.
Veröffentlicht: (2025) -
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025) -
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)