SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
Fuente:
arXiv
Saved in:
| Main Author: | Shukla, Shikhar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecMemo: Speculative Decoding is in Your Pocket
by: Yildirim, Selin, et al.
Published: (2025)
by: Yildirim, Selin, et al.
Published: (2025)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Make Every Draft Count: Hidden State based Speculative Decoding
by: Chen, Yuetao, et al.
Published: (2026)
by: Chen, Yuetao, et al.
Published: (2026)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
by: Wu, Hang, et al.
Published: (2025)
by: Wu, Hang, et al.
Published: (2025)
VREM-FL: Mobility-Aware Computation-Scheduling Co-Design for Vehicular Federated Learning
by: Ballotta, Luca, et al.
Published: (2023)
by: Ballotta, Luca, et al.
Published: (2023)
When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
by: Wei, Jiankun, et al.
Published: (2024)
by: Wei, Jiankun, et al.
Published: (2024)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
by: Timor, Nadav, et al.
Published: (2024)
by: Timor, Nadav, et al.
Published: (2024)
Utility-Driven Speculative Decoding for Mixture-of-Experts
by: Saxena, Anish, et al.
Published: (2025)
by: Saxena, Anish, et al.
Published: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
CA-AFP: Cluster-Aware Adaptive Federated Pruning
by: Jha, Om Govind, et al.
Published: (2026)
by: Jha, Om Govind, et al.
Published: (2026)
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum
by: De Silva, Suvi, et al.
Published: (2026)
by: De Silva, Suvi, et al.
Published: (2026)
Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
by: Taufique, Zain, et al.
Published: (2023)
by: Taufique, Zain, et al.
Published: (2023)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
by: Hu, Xinyi, et al.
Published: (2026)
by: Hu, Xinyi, et al.
Published: (2026)
AMV-L: Lifecycle-Managed Agent Memory for Tail-Latency Control in Long-Running LLM Systems
by: Bamidele, Emmanuel
Published: (2026)
by: Bamidele, Emmanuel
Published: (2026)
Soar: Design and Deployment of A Smart Roadside Infrastructure System for Autonomous Driving
by: Shi, Shuyao, et al.
Published: (2024)
by: Shi, Shuyao, et al.
Published: (2024)
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
by: Wang, Taiyi, et al.
Published: (2024)
by: Wang, Taiyi, et al.
Published: (2024)
Improving accuracy and convergence of federated learning edge computing methods for generalized DER forecasting applications in power grid
by: Nair, Vineet Jagadeesan, et al.
Published: (2024)
by: Nair, Vineet Jagadeesan, et al.
Published: (2024)
Evaluation of a Foundational Model and Stochastic Models for Forecasting Sporadic or Spiky Production Outages of High-Performance Machine Learning Services
by: Yim, Keun Soo
Published: (2025)
by: Yim, Keun Soo
Published: (2025)
Fully Decentralized Joint Learning of Personalized Models and Collaboration Graphs
by: Zantedeschi, Valentina, et al.
Published: (2019)
by: Zantedeschi, Valentina, et al.
Published: (2019)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
by: Guo, Yipin, et al.
Published: (2026)
by: Guo, Yipin, et al.
Published: (2026)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
by: Patel, Ishan, et al.
Published: (2026)
by: Patel, Ishan, et al.
Published: (2026)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
by: McDanel, Bradley
Published: (2024)
by: McDanel, Bradley
Published: (2024)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
by: Li, Weiqing, et al.
Published: (2025)
by: Li, Weiqing, et al.
Published: (2025)
Multi-Bin Batching for Increasing LLM Inference Throughput
by: Guldogan, Ozgur, et al.
Published: (2024)
by: Guldogan, Ozgur, et al.
Published: (2024)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
Lifelong Federated Reinforcement Learning: A Learning Architecture for Navigation in Cloud Robotic Systems
by: Liu, Boyi, et al.
Published: (2019)
by: Liu, Boyi, et al.
Published: (2019)
Equilibrium in the Computing Continuum through Active Inference
by: Sedlak, Boris, et al.
Published: (2023)
by: Sedlak, Boris, et al.
Published: (2023)
Federated reinforcement learning for robot motion planning with zero-shot generalization
by: Yuan, Zhenyuan, et al.
Published: (2024)
by: Yuan, Zhenyuan, et al.
Published: (2024)
The Smart Buildings Control Suite: A Diverse Open Source Benchmark to Evaluate and Scale HVAC Control Policies for Sustainability
by: Goldfeder, Judah, et al.
Published: (2024)
by: Goldfeder, Judah, et al.
Published: (2024)
Sustainability of Data Center Digital Twins with Reinforcement Learning
by: Sarkar, Soumyendu, et al.
Published: (2024)
by: Sarkar, Soumyendu, et al.
Published: (2024)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025)
by: Xu, Jiaming, et al.
Published: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Similar Items
-
SpecMemo: Speculative Decoding is in Your Pocket
by: Yildirim, Selin, et al.
Published: (2025) -
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024) -
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026) -
Make Every Draft Count: Hidden State based Speculative Decoding
by: Chen, Yuetao, et al.
Published: (2026) -
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)