Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Zixiao, Zeng, Wen, Fu, Tianyu, Liu, Tengxuan, Sun, Yizhou, Hong, Ke, Yang, Xinhao, Liu, Chengchun, Li, Yan, Zhang, Quanlu, Dai, Guohao, Zhu, Zhenhua, Wang, Yu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
An Interpretable Latency Model for Speculative Decoding in LLM Serving
por: Kong, Linghao, et al.
Publicado: (2026)
por: Kong, Linghao, et al.
Publicado: (2026)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
por: Huang, Zixiao, et al.
Publicado: (2025)
por: Huang, Zixiao, et al.
Publicado: (2025)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
por: Zhang, Yijia, et al.
Publicado: (2024)
por: Zhang, Yijia, et al.
Publicado: (2024)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
por: Liu, Xiaoxuan, et al.
Publicado: (2024)
por: Liu, Xiaoxuan, et al.
Publicado: (2024)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
por: Zhou, Fang, et al.
Publicado: (2026)
por: Zhou, Fang, et al.
Publicado: (2026)
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
por: Yin, Wangsong, et al.
Publicado: (2025)
por: Yin, Wangsong, et al.
Publicado: (2025)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
por: Jeffery, Andrew, et al.
Publicado: (2024)
por: Jeffery, Andrew, et al.
Publicado: (2024)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
por: Wang, Haoxin, et al.
Publicado: (2025)
por: Wang, Haoxin, et al.
Publicado: (2025)
Stability and Optimization of Speculative Queueing Networks
por: Anselmi, Jonatha, et al.
Publicado: (2021)
por: Anselmi, Jonatha, et al.
Publicado: (2021)
Two Criteria for Performance Analysis of Optimization Algorithms
por: Jing, Yunpeng, et al.
Publicado: (2024)
por: Jing, Yunpeng, et al.
Publicado: (2024)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
por: Fu, Tianyu, et al.
Publicado: (2025)
por: Fu, Tianyu, et al.
Publicado: (2025)
EDAN: Towards Understanding Memory Parallelism and Latency Sensitivity in HPC
por: Shen, Siyuan, et al.
Publicado: (2025)
por: Shen, Siyuan, et al.
Publicado: (2025)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
por: Ahmadvand, Hossein, et al.
Publicado: (2025)
por: Ahmadvand, Hossein, et al.
Publicado: (2025)
On Latency Predictors for Neural Architecture Search
por: Akhauri, Yash, et al.
Publicado: (2024)
por: Akhauri, Yash, et al.
Publicado: (2024)
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
por: Chen, Mingkai, et al.
Publicado: (2024)
por: Chen, Mingkai, et al.
Publicado: (2024)
Steering Pretrained Drafters during Speculative Decoding
por: Berdoz, Frédéric, et al.
Publicado: (2025)
por: Berdoz, Frédéric, et al.
Publicado: (2025)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
por: Liu, Qunyou, et al.
Publicado: (2025)
por: Liu, Qunyou, et al.
Publicado: (2025)
A Latency-Constrained, Gated Recurrent Unit (GRU) Implementation in the Versal AI Engine
por: Sapkas, M., et al.
Publicado: (2025)
por: Sapkas, M., et al.
Publicado: (2025)
PM2Lat: Highly Accurate and Generalized Prediction of DNN Execution Latency on GPUs
por: Le, Truong-Thanh, et al.
Publicado: (2026)
por: Le, Truong-Thanh, et al.
Publicado: (2026)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
por: Fu, Zizhuo, et al.
Publicado: (2025)
por: Fu, Zizhuo, et al.
Publicado: (2025)
Motion-to-Motion Latency Measurement Framework for Connected and Autonomous Vehicle Teleoperation
por: Provost, François, et al.
Publicado: (2025)
por: Provost, François, et al.
Publicado: (2025)
An Experimental Study of Low-Latency Video Streaming over 5G
por: Khan, Imran, et al.
Publicado: (2024)
por: Khan, Imran, et al.
Publicado: (2024)
Reducing Waiting Time for Medical Tourists Through Hybrid Agent-Based and Discrete-Event Simulation: A Hospital Case Study
por: Baghi, Melika, et al.
Publicado: (2026)
por: Baghi, Melika, et al.
Publicado: (2026)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
por: Sharma, Vinamra, et al.
Publicado: (2026)
por: Sharma, Vinamra, et al.
Publicado: (2026)
Latency Based Tiling
por: Cashman, Jack
Publicado: (2025)
por: Cashman, Jack
Publicado: (2025)
PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor
por: Zhao, Xinlong, et al.
Publicado: (2025)
por: Zhao, Xinlong, et al.
Publicado: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
por: Tao, Yiheng, et al.
Publicado: (2025)
por: Tao, Yiheng, et al.
Publicado: (2025)
Pinching-Antenna Systems For Indoor Immersive Communications: A 3D-Modeling Based Performance Analysis
por: Wang, Yulei, et al.
Publicado: (2025)
por: Wang, Yulei, et al.
Publicado: (2025)
Bayesian Hierarchical Models for Quantitative Estimates for Performance metrics applied to Saddle Search Algorithms
por: Goswami, Rohit
Publicado: (2025)
por: Goswami, Rohit
Publicado: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
por: Dong, Ximing, et al.
Publicado: (2026)
por: Dong, Ximing, et al.
Publicado: (2026)
Analysis and Evaluation of Using Microsecond-Latency Memory for In-Memory Indices and Caches in SSD-Based Key-Value Stores
por: Bando, Yosuke, et al.
Publicado: (2025)
por: Bando, Yosuke, et al.
Publicado: (2025)
On the Design of Capacity-Achieving Distributions for Discrete-Time Poisson Channel with Low-Precision ADCs
por: Li, Qianqian, et al.
Publicado: (2025)
por: Li, Qianqian, et al.
Publicado: (2025)
Approximation Algorithms for Minimizing Congestion in Demand-Aware Networks
por: Dai, Wenkai, et al.
Publicado: (2024)
por: Dai, Wenkai, et al.
Publicado: (2024)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
por: Fu, Tianyu, et al.
Publicado: (2025)
por: Fu, Tianyu, et al.
Publicado: (2025)
Towards A Flexible Accuracy-Oriented Deep Learning Module Inference Latency Prediction Framework for Adaptive Optimization Algorithms
por: Shen, Jingran, et al.
Publicado: (2023)
por: Shen, Jingran, et al.
Publicado: (2023)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
por: Zhang, Yaozheng, et al.
Publicado: (2025)
por: Zhang, Yaozheng, et al.
Publicado: (2025)
On Combining Two Server Control Policies for Energy Efficiency
por: Dai, Jingze, et al.
Publicado: (2025)
por: Dai, Jingze, et al.
Publicado: (2025)
Shortest-Path FFT: Optimal SIMD Instruction Scheduling via Graph Search
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Ejemplares similares
-
An Interpretable Latency Model for Speculative Decoding in LLM Serving
por: Kong, Linghao, et al.
Publicado: (2026) -
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
por: Huang, Zixiao, et al.
Publicado: (2025) -
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
por: Zhang, Yijia, et al.
Publicado: (2024) -
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
por: Liu, Xiaoxuan, et al.
Publicado: (2024) -
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
por: Zhou, Fang, et al.
Publicado: (2026)