When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Tianyu, Shen, Yuhao, Hu, Xinyi, Zhang, Baolin, Zhang, Hengxin, Dai, Jun, Zhang, Jun, Ge, Shuang, Chen, Lei, Li, Yue, Wan, MingCheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
von: Hao, Jitai, et al.
Veröffentlicht: (2026)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMs
von: Qi, Yanlin, et al.
Veröffentlicht: (2026)
von: Qi, Yanlin, et al.
Veröffentlicht: (2026)
Competitive Non-Clairvoyant KV-Cache Scheduling for LLM Inference
von: Feng, Yiding, et al.
Veröffentlicht: (2026)
von: Feng, Yiding, et al.
Veröffentlicht: (2026)
Moment-KV: Momentum-Based Decode-Time KV Cache Compression for Long Generation
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
von: Jana, Soumyadeep, et al.
Veröffentlicht: (2026)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
von: Fang, Min, et al.
Veröffentlicht: (2025)
von: Fang, Min, et al.
Veröffentlicht: (2025)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
von: Wan, Zhongwei, et al.
Veröffentlicht: (2025)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
von: Wu, Wenbo, et al.
Veröffentlicht: (2025)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
von: Eldenk, Doğaç, et al.
Veröffentlicht: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
von: Cai, Zefan, et al.
Veröffentlicht: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
von: Liu, Tianyu, et al.
Veröffentlicht: (2026)
When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
von: Wei, Jiankun, et al.
Veröffentlicht: (2024)
von: Wei, Jiankun, et al.
Veröffentlicht: (2024)
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
CacheClip: Accelerating RAG with Effective KV Cache Reuse
von: Yang, Bin, et al.
Veröffentlicht: (2025)
von: Yang, Bin, et al.
Veröffentlicht: (2025)
LagKV: Lag-Relative Information of the KV Cache Tells Which Tokens Are Important
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
von: Liang, Manlai, et al.
Veröffentlicht: (2025)
HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference
von: Zeng, Bowen, et al.
Veröffentlicht: (2026)
von: Zeng, Bowen, et al.
Veröffentlicht: (2026)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
von: Guo, Jinyu, et al.
Veröffentlicht: (2026)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
von: Zhou, Xiabin, et al.
Veröffentlicht: (2024)
LongFlow: Efficient KV Cache Compression for Reasoning Models
von: Su, Yi, et al.
Veröffentlicht: (2026)
von: Su, Yi, et al.
Veröffentlicht: (2026)
Scaling Speculative Decoding with Lookahead Reasoning
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
von: Hu, Jie, et al.
Veröffentlicht: (2025)
von: Hu, Jie, et al.
Veröffentlicht: (2025)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
von: Sun, Qiheng, et al.
Veröffentlicht: (2025)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
von: Li, Weizhuo, et al.
Veröffentlicht: (2024)
von: Li, Weizhuo, et al.
Veröffentlicht: (2024)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
von: Ge, Suyu, et al.
Veröffentlicht: (2023)
von: Ge, Suyu, et al.
Veröffentlicht: (2023)
When Drafts Evolve: Speculative Decoding Meets Online Learning
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
von: Jo, Dongwon, et al.
Veröffentlicht: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025)
von: Du, Dayou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
von: Hu, Xinyi, et al.
Veröffentlicht: (2026) -
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026) -
DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity
von: Hao, Jitai, et al.
Veröffentlicht: (2026) -
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)