An Interpretable Latency Model for Speculative Decoding in LLM Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Linghao, Flynn, Megan, Peng, Michael, Shavit, Nir, Kurtz, Mark, Marques, Alexandre |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
Negative Pre-activations Differentiate Syntax
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
Fairness in Serving Large Language Models
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
Expand Neurons, Not Parameters
by: Kong, Linghao, et al.
Published: (2025)
by: Kong, Linghao, et al.
Published: (2025)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
by: Hendria, Willy Fitra
Published: (2026)
by: Hendria, Willy Fitra
Published: (2026)
Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems
by: Zhao, Yushang, et al.
Published: (2025)
by: Zhao, Yushang, et al.
Published: (2025)
Toy Combinatorial Interpretability Models Reveal Lottery Tickets in Early Feature Space
by: Bebchuk, Alon, et al.
Published: (2026)
by: Bebchuk, Alon, et al.
Published: (2026)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
by: Iglovikov, Vladimir, et al.
Published: (2026)
by: Iglovikov, Vladimir, et al.
Published: (2026)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
Wasserstein Distances, Neuronal Entanglement, and Sparsity
by: Sawmya, Shashata, et al.
Published: (2024)
by: Sawmya, Shashata, et al.
Published: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
by: Hu, Xiannan, et al.
Published: (2025)
by: Hu, Xiannan, et al.
Published: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
by: Barad, Haim, et al.
Published: (2023)
by: Barad, Haim, et al.
Published: (2023)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
An Inquiry into Datacenter TCO for LLM Inference with FP8
by: Kim, Jiwoo, et al.
Published: (2025)
by: Kim, Jiwoo, et al.
Published: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
by: Yu, Shan, et al.
Published: (2025)
by: Yu, Shan, et al.
Published: (2025)
CEBench: A Benchmarking Toolkit for the Cost-Effectiveness of LLM Pipelines
by: Sun, Wenbo, et al.
Published: (2024)
by: Sun, Wenbo, et al.
Published: (2024)
Development and Comparative Evaluation of Three Artificial Intelligence Models (NLP, LLM, JEPA) for Predicting Triage in Emergency Departments: A 7-Month Retrospective Proof-of-Concept
by: Lansiaux, Edouard, et al.
Published: (2025)
by: Lansiaux, Edouard, et al.
Published: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
by: Chu, Kexin, et al.
Published: (2026)
by: Chu, Kexin, et al.
Published: (2026)
Plug-and-Play Performance Estimation for LLM Services without Relying on Labeled Data
by: Wang, Can, et al.
Published: (2024)
by: Wang, Can, et al.
Published: (2024)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
by: Ma, Xinyue, et al.
Published: (2026)
by: Ma, Xinyue, et al.
Published: (2026)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
Prioritizing Latency with Profit: A DRL-Based Admission Control for 5G Network Slices
by: Chakraborty, Proggya, et al.
Published: (2025)
by: Chakraborty, Proggya, et al.
Published: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
A Study on Inference Latency for Vision Transformers on Mobile Devices
by: Li, Zhuojin, et al.
Published: (2025)
by: Li, Zhuojin, et al.
Published: (2025)
Ensuring Reliability of Curated EHR-Derived Data: The Validation of Accuracy for LLM/ML-Extracted Information and Data (VALID) Framework
by: Estevez, Melissa, et al.
Published: (2025)
by: Estevez, Melissa, et al.
Published: (2025)
ALISE: Accelerating Large Language Model Serving with Speculative Scheduling
by: Zhao, Youpeng, et al.
Published: (2024)
by: Zhao, Youpeng, et al.
Published: (2024)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
by: Atmer, Hannah, et al.
Published: (2025)
by: Atmer, Hannah, et al.
Published: (2025)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
by: Liu, Xiaoxuan, et al.
Published: (2024)
by: Liu, Xiaoxuan, et al.
Published: (2024)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
by: Abraham, Ashley N., et al.
Published: (2026)
by: Abraham, Ashley N., et al.
Published: (2026)
Towards A Flexible Accuracy-Oriented Deep Learning Module Inference Latency Prediction Framework for Adaptive Optimization Algorithms
by: Shen, Jingran, et al.
Published: (2023)
by: Shen, Jingran, et al.
Published: (2023)
An Autotuning-based Optimization Framework for Mixed-kernel SVM Classifications in Smart Pixel Datasets and Heterojunction Transistors
by: Wu, Xingfu, et al.
Published: (2024)
by: Wu, Xingfu, et al.
Published: (2024)
Benchmarking GPUs on SVBRDF Extractor Model
by: Kandel, Narayan, et al.
Published: (2023)
by: Kandel, Narayan, et al.
Published: (2023)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
by: Zhang, Hang, et al.
Published: (2025)
by: Zhang, Hang, et al.
Published: (2025)
On Latency Predictors for Neural Architecture Search
by: Akhauri, Yash, et al.
Published: (2024)
by: Akhauri, Yash, et al.
Published: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
Similar Items
-
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025) -
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025) -
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025) -
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026) -
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)