Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hooper, Coleman, Kang, Minwoo, Moon, Suhong, Lee, Nicholas, Wen, Eric, Wawrzynek, John, Mahoney, Michael W., Shao, Yakun Sophia, Gholami, Amir, Keutzer, Kurt |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SPEED: Speculative Pipelined Execution for Efficient Decoding
par: Hooper, Coleman, et autres
Publié: (2023)
par: Hooper, Coleman, et autres
Publié: (2023)
TinyAgent: Function Calling at the Edge
par: Erdogan, Lutfi Eren, et autres
Publié: (2024)
par: Erdogan, Lutfi Eren, et autres
Publié: (2024)
An LLM Compiler for Parallel Function Calling
par: Kim, Sehoon, et autres
Publié: (2023)
par: Kim, Sehoon, et autres
Publié: (2023)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
par: Hooper, Coleman, et autres
Publié: (2024)
par: Hooper, Coleman, et autres
Publié: (2024)
ETS: Efficient Tree Search for Inference-Time Scaling
par: Hooper, Coleman, et autres
Publié: (2025)
par: Hooper, Coleman, et autres
Publié: (2025)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
par: Maheswaran, Monishwaran, et autres
Publié: (2025)
par: Maheswaran, Monishwaran, et autres
Publié: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
par: Tiwari, Rishabh, et autres
Publié: (2025)
par: Tiwari, Rishabh, et autres
Publié: (2025)
Multipole Attention for Efficient Long Context Reasoning
par: Hooper, Coleman, et autres
Publié: (2025)
par: Hooper, Coleman, et autres
Publié: (2025)
AI and Memory Wall
par: Gholami, Amir, et autres
Publié: (2024)
par: Gholami, Amir, et autres
Publié: (2024)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
par: Erdogan, Lutfi Eren, et autres
Publié: (2025)
par: Erdogan, Lutfi Eren, et autres
Publié: (2025)
Efficient and Scalable Estimation of Tool Representations in Vector Space
par: Moon, Suhong, et autres
Publié: (2024)
par: Moon, Suhong, et autres
Publié: (2024)
Ghost Tool Calls: Issue-Time Privacy for Speculative Agent Tools
par: Mohammadi, Bardia, et autres
Publié: (2026)
par: Mohammadi, Bardia, et autres
Publié: (2026)
SqueezeLLM: Dense-and-Sparse Quantization
par: Kim, Sehoon, et autres
Publié: (2023)
par: Kim, Sehoon, et autres
Publié: (2023)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
par: Kim, Minseo, et autres
Publié: (2025)
par: Kim, Minseo, et autres
Publié: (2025)
Agentic Test-Time Scaling for WebAgents
par: Lee, Nicholas, et autres
Publié: (2026)
par: Lee, Nicholas, et autres
Publié: (2026)
DEMOTIC: A Differentiable Sampler for Multi-Level Digital Circuits
par: Ardakani, Arash, et autres
Publié: (2025)
par: Ardakani, Arash, et autres
Publié: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
par: Hooper, Coleman, et autres
Publié: (2024)
par: Hooper, Coleman, et autres
Publié: (2024)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
par: Tomar, Aditya, et autres
Publié: (2025)
par: Tomar, Aditya, et autres
Publié: (2025)
SciML Agents: Write the Solver, Not the Solution
par: Gaonkar, Saarth, et autres
Publié: (2025)
par: Gaonkar, Saarth, et autres
Publié: (2025)
CDLM: Consistency Diffusion Language Models For Faster Sampling
par: Kim, Minseo, et autres
Publié: (2025)
par: Kim, Minseo, et autres
Publié: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
par: Hooper, Coleman, et autres
Publié: (2025)
par: Hooper, Coleman, et autres
Publié: (2025)
Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans
par: Moon, Suhong, et autres
Publié: (2026)
par: Moon, Suhong, et autres
Publié: (2026)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
par: Xi, Haocheng, et autres
Publié: (2026)
par: Xi, Haocheng, et autres
Publié: (2026)
Learned Best-Effort LLM Serving
par: Jha, Siddharth, et autres
Publié: (2024)
par: Jha, Siddharth, et autres
Publié: (2024)
Dynamic Speculative Agent Planning
par: Guan, Yilin, et autres
Publié: (2025)
par: Guan, Yilin, et autres
Publié: (2025)
Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors
par: Suh, Joseph, et autres
Publié: (2024)
par: Suh, Joseph, et autres
Publié: (2024)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
par: Subramanian, Shashank, et autres
Publié: (2023)
par: Subramanian, Shashank, et autres
Publié: (2023)
Optimizing Agentic Language Model Inference via Speculative Tool Calls
par: Nichols, Daniel, et autres
Publié: (2025)
par: Nichols, Daniel, et autres
Publié: (2025)
Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions
par: Suh, Joseph, et autres
Publié: (2025)
par: Suh, Joseph, et autres
Publié: (2025)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
par: Cemri, Mert, et autres
Publié: (2025)
par: Cemri, Mert, et autres
Publié: (2025)
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
par: Xia, Heming, et autres
Publié: (2026)
par: Xia, Heming, et autres
Publié: (2026)
Asynchronous Tool Usage for Real-Time Agents
par: Ginart, Antonio A., et autres
Publié: (2024)
par: Ginart, Antonio A., et autres
Publié: (2024)
Residual Context Diffusion Language Models
par: Hu, Yuezhou, et autres
Publié: (2026)
par: Hu, Yuezhou, et autres
Publié: (2026)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
par: Lee, Nicholas, et autres
Publié: (2024)
par: Lee, Nicholas, et autres
Publié: (2024)
Speculative Speculative Decoding
par: Kumar, Tanishq, et autres
Publié: (2026)
par: Kumar, Tanishq, et autres
Publié: (2026)
High-Throughput SAT Sampling
par: Ardakani, Arash, et autres
Publié: (2025)
par: Ardakani, Arash, et autres
Publié: (2025)
Characterizing Prompt Compression Methods for Long Context Inference
par: Jha, Siddharth, et autres
Publié: (2024)
par: Jha, Siddharth, et autres
Publié: (2024)
SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion
par: Ma, George, et autres
Publié: (2025)
par: Ma, George, et autres
Publié: (2025)
Speculations on Uncertainty and Humane Algorithms
par: Gray, Nicholas
Publié: (2024)
par: Gray, Nicholas
Publié: (2024)
MeNTi: Bridging Medical Calculator and LLM Agent with Nested Tool Calling
par: Zhu, Yakun, et autres
Publié: (2024)
par: Zhu, Yakun, et autres
Publié: (2024)
Documents similaires
-
SPEED: Speculative Pipelined Execution for Efficient Decoding
par: Hooper, Coleman, et autres
Publié: (2023) -
TinyAgent: Function Calling at the Edge
par: Erdogan, Lutfi Eren, et autres
Publié: (2024) -
An LLM Compiler for Parallel Function Calling
par: Kim, Sehoon, et autres
Publié: (2023) -
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
par: Hooper, Coleman, et autres
Publié: (2024) -
ETS: Efficient Tree Search for Inference-Time Scaling
par: Hooper, Coleman, et autres
Publié: (2025)