SimpleTool: Parallel Decoding for Real-Time LLM Function Calling
Fuente:
arXiv
Guardado en:
| Autores principales: | Shi, Xiaoxin, Wan, Jiaxin, Dong, Linkang, Jiang, Wei, Liu, Yue, Huang, Zengfeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An LLM Compiler for Parallel Function Calling
por: Kim, Sehoon, et al.
Publicado: (2023)
por: Kim, Sehoon, et al.
Publicado: (2023)
ToolACE: Winning the Points of LLM Function Calling
por: Liu, Weiwen, et al.
Publicado: (2024)
por: Liu, Weiwen, et al.
Publicado: (2024)
Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams
por: Sun, Tianda, et al.
Publicado: (2026)
por: Sun, Tianda, et al.
Publicado: (2026)
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
por: Wang, Zezhong, et al.
Publicado: (2024)
por: Wang, Zezhong, et al.
Publicado: (2024)
LLM Agents Already Know When to Call Tools -- Even Without Reasoning
por: Sun, Chung-En, et al.
Publicado: (2026)
por: Sun, Chung-En, et al.
Publicado: (2026)
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
por: Xia, Heming, et al.
Publicado: (2026)
por: Xia, Heming, et al.
Publicado: (2026)
TOOL-ED: Enhancing Empathetic Response Generation with the Tool Calling Capability of LLM
por: Cao, Huiying, et al.
Publicado: (2024)
por: Cao, Huiying, et al.
Publicado: (2024)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
por: Xu, Haoyuan, et al.
Publicado: (2026)
por: Xu, Haoyuan, et al.
Publicado: (2026)
Uncertainty Quantification for LLM Function-Calling
por: Ye, Zihuiwen, et al.
Publicado: (2026)
por: Ye, Zihuiwen, et al.
Publicado: (2026)
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
por: Liu, Ziyang
Publicado: (2026)
por: Liu, Ziyang
Publicado: (2026)
RelayLLM: Efficient Reasoning via Collaborative Decoding
por: Huang, Chengsong, et al.
Publicado: (2026)
por: Huang, Chengsong, et al.
Publicado: (2026)
Asynchronous LLM Function Calling
por: Gim, In, et al.
Publicado: (2024)
por: Gim, In, et al.
Publicado: (2024)
When2Call: When (not) to Call Tools
por: Ross, Hayley, et al.
Publicado: (2025)
por: Ross, Hayley, et al.
Publicado: (2025)
A Frustratingly Simple Decoding Method for Neural Text Generation
por: Yang, Haoran, et al.
Publicado: (2023)
por: Yang, Haoran, et al.
Publicado: (2023)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
por: Zhang, Kangning, et al.
Publicado: (2025)
por: Zhang, Kangning, et al.
Publicado: (2025)
MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling
por: Zhu, Yakun, et al.
Publicado: (2024)
por: Zhu, Yakun, et al.
Publicado: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
por: Xiao, Zilin, et al.
Publicado: (2024)
por: Xiao, Zilin, et al.
Publicado: (2024)
HAF-RM: A Hybrid Alignment Framework for Reward Model Training
por: Liu, Shujun, et al.
Publicado: (2024)
por: Liu, Shujun, et al.
Publicado: (2024)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
por: Wu, Chengyue, et al.
Publicado: (2025)
por: Wu, Chengyue, et al.
Publicado: (2025)
Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
por: Song, Yuerong, et al.
Publicado: (2025)
por: Song, Yuerong, et al.
Publicado: (2025)
PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity Recognition
por: Lu, Jinghui, et al.
Publicado: (2024)
por: Lu, Jinghui, et al.
Publicado: (2024)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
por: Zhong, Shuzhang, et al.
Publicado: (2024)
por: Zhong, Shuzhang, et al.
Publicado: (2024)
MARSAD: A Multi-Functional Tool for Real-Time Social Media Analysis
por: Biswas, Md. Rafiul, et al.
Publicado: (2025)
por: Biswas, Md. Rafiul, et al.
Publicado: (2025)
CallNavi, A Challenge and Empirical Study on LLM Function Calling and Routing
por: Song, Yewei, et al.
Publicado: (2025)
por: Song, Yewei, et al.
Publicado: (2025)
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
por: Liu, Mingdao, et al.
Publicado: (2024)
por: Liu, Mingdao, et al.
Publicado: (2024)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
por: Lu, Yuxuan, et al.
Publicado: (2026)
por: Lu, Yuxuan, et al.
Publicado: (2026)
Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
por: Liu, Xiaoran, et al.
Publicado: (2025)
por: Liu, Xiaoran, et al.
Publicado: (2025)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
por: Mohammadi, Bardia, et al.
Publicado: (2026)
por: Mohammadi, Bardia, et al.
Publicado: (2026)
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
por: Lumer, Elias, et al.
Publicado: (2025)
por: Lumer, Elias, et al.
Publicado: (2025)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
por: Xu, Chenkai, et al.
Publicado: (2025)
por: Xu, Chenkai, et al.
Publicado: (2025)
dParallel: Learnable Parallel Decoding for dLLMs
por: Chen, Zigeng, et al.
Publicado: (2025)
por: Chen, Zigeng, et al.
Publicado: (2025)
DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents
por: Zhang, Junshuo, et al.
Publicado: (2026)
por: Zhang, Junshuo, et al.
Publicado: (2026)
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
por: Laskar, Md Tahmid Rahman, et al.
Publicado: (2026)
por: Laskar, Md Tahmid Rahman, et al.
Publicado: (2026)
Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing
por: Zheng, Tong, et al.
Publicado: (2026)
por: Zheng, Tong, et al.
Publicado: (2026)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
por: Liu, Tianyu, et al.
Publicado: (2024)
por: Liu, Tianyu, et al.
Publicado: (2024)
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
por: Du, Yu, et al.
Publicado: (2024)
por: Du, Yu, et al.
Publicado: (2024)
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit
por: Wang, Kangyu, et al.
Publicado: (2025)
por: Wang, Kangyu, et al.
Publicado: (2025)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
por: Svirschevski, Ruslan, et al.
Publicado: (2024)
por: Svirschevski, Ruslan, et al.
Publicado: (2024)
Ejemplares similares
-
An LLM Compiler for Parallel Function Calling
por: Kim, Sehoon, et al.
Publicado: (2023) -
ToolACE: Winning the Points of LLM Function Calling
por: Liu, Weiwen, et al.
Publicado: (2024) -
Tool-Call Dependency Structure is Linearly Decodable in LLM Agent Residual Streams
por: Sun, Tianda, et al.
Publicado: (2026) -
AdaDecode: Accelerating LLM Decoding with Adaptive Layer Parallelism
por: Wei, Zhepei, et al.
Publicado: (2025) -
An LLM-Tool Compiler for Fused Parallel Function Calling
por: Singh, Simranjit, et al.
Publicado: (2024)