Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution
Fuente:
arXiv
Saved in:
| Main Authors: | Sui, Yifan, Zhao, Han, Ma, Rui, He, Zhiyuan, Wang, Hao, Li, Jianxun, Yang, Yuqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
by: Song, Yanfei
Published: (2026)
by: Song, Yanfei
Published: (2026)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024)
by: Wang, Siqi, et al.
Published: (2024)
Distributed Speculative Execution for Resilient Cloud Applications
by: Li, Tianyu, et al.
Published: (2024)
by: Li, Tianyu, et al.
Published: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025)
by: Sui, Yifan, et al.
Published: (2025)
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
by: He, Zijian, et al.
Published: (2026)
by: He, Zijian, et al.
Published: (2026)
A Unified Programming Model for Heterogeneous Computing with CPU and Accelerator Technologies
by: Xiong, Yuqing
Published: (2022)
by: Xiong, Yuqing
Published: (2022)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)
by: Dai, Yuntao, et al.
Published: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
by: Chen, Fahao, et al.
Published: (2025)
by: Chen, Fahao, et al.
Published: (2025)
LAAFD: LLM-based Agents for Accelerated FPGA Design
by: Moraru, Maxim, et al.
Published: (2026)
by: Moraru, Maxim, et al.
Published: (2026)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
Data-Locality-Aware Task Assignment and Scheduling for Distributed Job Executions
by: Zhao, Hailiang, et al.
Published: (2024)
by: Zhao, Hailiang, et al.
Published: (2024)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
by: Zheng, Ce, et al.
Published: (2026)
by: Zheng, Ce, et al.
Published: (2026)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
by: Chen, Wenyan, et al.
Published: (2026)
by: Chen, Wenyan, et al.
Published: (2026)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
by: Yuan, Yitao, et al.
Published: (2025)
by: Yuan, Yitao, et al.
Published: (2025)
Dependency-Aware Execution Mechanism in Hyperledger Fabric Architecture
by: Kaul, Sanyam, et al.
Published: (2025)
by: Kaul, Sanyam, et al.
Published: (2025)
Exploring the Potential of Carbon-Aware Execution for Scientific Workflows
by: West, Kathleen, et al.
Published: (2025)
by: West, Kathleen, et al.
Published: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
PipeLLM: Fast and Confidential Large Language Model Services with Speculative Pipelined Encryption
by: Tan, Yifan, et al.
Published: (2024)
by: Tan, Yifan, et al.
Published: (2024)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Object Proxy Patterns for Accelerating Distributed Applications
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
Zenix: Efficient Execution of Bulky Serverless Applications
by: Guo, Zhiyuan, et al.
Published: (2022)
by: Guo, Zhiyuan, et al.
Published: (2022)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
by: Qian, Daniel, et al.
Published: (2026)
by: Qian, Daniel, et al.
Published: (2026)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
by: Tran, Phuong, et al.
Published: (2025)
by: Tran, Phuong, et al.
Published: (2025)
A Systematic Evaluation of the Potential of Carbon-Aware Execution for Scientific Workflows
by: West, Kathleen, et al.
Published: (2025)
by: West, Kathleen, et al.
Published: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
by: Ma, Mulei, et al.
Published: (2025)
by: Ma, Mulei, et al.
Published: (2025)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
by: Zhao, Zhan, et al.
Published: (2026)
by: Zhao, Zhan, et al.
Published: (2026)
Hierarchical Observe-Orient-Decide-Act Enabled UAV Swarms in Uncertain Environments: Frameworks, Potentials, and Challenges
by: Jia, Ziye, et al.
Published: (2026)
by: Jia, Ziye, et al.
Published: (2026)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
by: Szafarczyk, Robert, et al.
Published: (2025)
by: Szafarczyk, Robert, et al.
Published: (2025)
PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernel
by: Yi, Jinjun, et al.
Published: (2025)
by: Yi, Jinjun, et al.
Published: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
by: Zhang, Han, et al.
Published: (2026)
by: Zhang, Han, et al.
Published: (2026)
A Comparative Evaluation of Automated Analysis Tools for Solidity Smart Contracts
by: Wei, Zhiyuan, et al.
Published: (2023)
by: Wei, Zhiyuan, et al.
Published: (2023)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
by: Kubo, Tatsuya, et al.
Published: (2025)
by: Kubo, Tatsuya, et al.
Published: (2025)
Similar Items
-
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
by: Song, Yanfei
Published: (2026) -
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024) -
Distributed Speculative Execution for Resilient Cloud Applications
by: Li, Tianyu, et al.
Published: (2024) -
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025) -
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
by: He, Zijian, et al.
Published: (2026)