An LLM Compiler for Parallel Function Calling
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Sehoon, Moon, Suhong, Tabrizi, Ryan, Lee, Nicholas, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TinyAgent: Function Calling at the Edge
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024)
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2025)
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2025)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
di: Lee, Nicholas, et al.
Pubblicazione: (2024)
di: Lee, Nicholas, et al.
Pubblicazione: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
di: Moon, Suhong, et al.
Pubblicazione: (2024)
di: Moon, Suhong, et al.
Pubblicazione: (2024)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
Squeezed Attention: Accelerating Long Context Length LLM Inference
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
Characterizing Prompt Compression Methods for Long Context Inference
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
Multipole Attention for Efficient Long Context Reasoning
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
SPEED: Speculative Pipelined Execution for Efficient Decoding
di: Hooper, Coleman, et al.
Pubblicazione: (2023)
di: Hooper, Coleman, et al.
Pubblicazione: (2023)
AI and Memory Wall
di: Gholami, Amir, et al.
Pubblicazione: (2024)
di: Gholami, Amir, et al.
Pubblicazione: (2024)
Agentic Test-Time Scaling for WebAgents
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
ETS: Efficient Tree Search for Inference-Time Scaling
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
Learned Best-Effort LLM Serving
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
di: Hooper, Coleman, et al.
Pubblicazione: (2026)
di: Hooper, Coleman, et al.
Pubblicazione: (2026)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
di: Nehrdich, Sebastian, et al.
Pubblicazione: (2026)
di: Nehrdich, Sebastian, et al.
Pubblicazione: (2026)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
di: Singh, Simranjit, et al.
Pubblicazione: (2024)
CDLM: Consistency Diffusion Language Models For Faster Sampling
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
Graph-Based Alternatives to LLMs for Human Simulation
di: Suh, Joseph, et al.
Pubblicazione: (2025)
di: Suh, Joseph, et al.
Pubblicazione: (2025)
SimpleTool: Parallel Decoding for Real-Time LLM Function Calling
di: Shi, Xiaoxin, et al.
Pubblicazione: (2026)
di: Shi, Xiaoxin, et al.
Pubblicazione: (2026)
Residual Context Diffusion Language Models
di: Hu, Yuezhou, et al.
Pubblicazione: (2026)
di: Hu, Yuezhou, et al.
Pubblicazione: (2026)
Asynchronous LLM Function Calling
di: Gim, In, et al.
Pubblicazione: (2024)
di: Gim, In, et al.
Pubblicazione: (2024)
Uncertainty Quantification for LLM Function-Calling
di: Ye, Zihuiwen, et al.
Pubblicazione: (2026)
di: Ye, Zihuiwen, et al.
Pubblicazione: (2026)
Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors
di: Suh, Joseph, et al.
Pubblicazione: (2024)
di: Suh, Joseph, et al.
Pubblicazione: (2024)
Simple and Effective Input Reformulations for Translation
di: Yu, Brian, et al.
Pubblicazione: (2023)
di: Yu, Brian, et al.
Pubblicazione: (2023)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
di: Nehrdich, Sebastian, et al.
Pubblicazione: (2024)
di: Nehrdich, Sebastian, et al.
Pubblicazione: (2024)
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions
di: Suh, Joseph, et al.
Pubblicazione: (2025)
di: Suh, Joseph, et al.
Pubblicazione: (2025)
Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans
di: Moon, Suhong, et al.
Pubblicazione: (2026)
di: Moon, Suhong, et al.
Pubblicazione: (2026)
Learning Adaptive Parallel Reasoning with Language Models
di: Pan, Jiayi, et al.
Pubblicazione: (2025)
di: Pan, Jiayi, et al.
Pubblicazione: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
CallNavi, A Challenge and Empirical Study on LLM Function Calling and Routing
di: Song, Yewei, et al.
Pubblicazione: (2025)
di: Song, Yewei, et al.
Pubblicazione: (2025)
Deep Binding of Language Model Virtual Personas: a Study on Approximating Political Partisan Misperceptions
di: Kang, Minwoo, et al.
Pubblicazione: (2025)
di: Kang, Minwoo, et al.
Pubblicazione: (2025)
FlowCompile: An Optimizing Compiler for Structured LLM Workflows
di: Li, Junyan, et al.
Pubblicazione: (2026)
di: Li, Junyan, et al.
Pubblicazione: (2026)
OMPar: Automatic Parallelization with AI-Driven Source-to-Source Compilation
di: Kadosh, Tal, et al.
Pubblicazione: (2024)
di: Kadosh, Tal, et al.
Pubblicazione: (2024)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
di: Subramanian, Shashank, et al.
Pubblicazione: (2023)
di: Subramanian, Shashank, et al.
Pubblicazione: (2023)
Cognitive Agent Compilation for Explicit Problem Solver Modeling
di: Moon, Hyeongdon, et al.
Pubblicazione: (2026)
di: Moon, Hyeongdon, et al.
Pubblicazione: (2026)
On the Robustness of Agentic Function Calling
di: Rabinovich, Ella, et al.
Pubblicazione: (2025)
di: Rabinovich, Ella, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TinyAgent: Function Calling at the Edge
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024) -
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2025) -
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
di: Lee, Nicholas, et al.
Pubblicazione: (2024) -
Efficient and Scalable Estimation of Tool Representations in Vector Space
di: Moon, Suhong, et al.
Pubblicazione: (2024) -
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)