Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Luohe, Li, Zuchao, Zhang, Lefei, Qi, Baoyuan, Liu, Guoming, Zhao, Hai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
by: Shi, Luohe, et al.
Published: (2025)
by: Shi, Luohe, et al.
Published: (2025)
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026)
by: Zhang, Zihong, et al.
Published: (2026)
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
by: Shi, Luohe, et al.
Published: (2024)
by: Shi, Luohe, et al.
Published: (2024)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
by: Yang, Haoqi, et al.
Published: (2025)
by: Yang, Haoqi, et al.
Published: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
by: Hu, Jiliang, et al.
Published: (2025)
by: Hu, Jiliang, et al.
Published: (2025)
Faster MoE LLM Inference for Extremely Large Models
by: Yang, Haoqi, et al.
Published: (2025)
by: Yang, Haoqi, et al.
Published: (2025)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
by: Song, Weixi, et al.
Published: (2023)
by: Song, Weixi, et al.
Published: (2023)
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba
by: Zou, Yuchen, et al.
Published: (2024)
by: Zou, Yuchen, et al.
Published: (2024)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
by: Zhang, Zihong, et al.
Published: (2025)
by: Zhang, Zihong, et al.
Published: (2025)
Batch Speculative Decoding Done Right
by: Zhang, Ranran Haoran, et al.
Published: (2025)
by: Zhang, Ranran Haoran, et al.
Published: (2025)
SirLLM: Streaming Infinite Retentive LLM
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
by: Ma, Xiangyu, et al.
Published: (2026)
by: Ma, Xiangyu, et al.
Published: (2026)
A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling Correction
by: Zeng, Xiangke, et al.
Published: (2024)
by: Zeng, Xiangke, et al.
Published: (2024)
ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language Models
by: Guo, Jiani, et al.
Published: (2025)
by: Guo, Jiani, et al.
Published: (2025)
GKT: A Novel Guidance-Based Knowledge Transfer Framework For Efficient Cloud-edge Collaboration LLM Deployment
by: Yao, Yao, et al.
Published: (2024)
by: Yao, Yao, et al.
Published: (2024)
VHASR: A Multimodal Speech Recognition System With Vision Hotwords
by: Hu, Jiliang, et al.
Published: (2024)
by: Hu, Jiliang, et al.
Published: (2024)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
by: Huang, Jianuo, et al.
Published: (2026)
by: Huang, Jianuo, et al.
Published: (2026)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference
by: Liu, Fuliang, et al.
Published: (2026)
by: Liu, Fuliang, et al.
Published: (2026)
Scaling Laws for Speculative Decoding
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Self Speculative Decoding for Diffusion Large Language Models
by: Gao, Yifeng, et al.
Published: (2025)
by: Gao, Yifeng, et al.
Published: (2025)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
by: Wu, Zhaoxuan, et al.
Published: (2025)
by: Wu, Zhaoxuan, et al.
Published: (2025)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
by: Tang, Bangsheng, et al.
Published: (2025)
by: Tang, Bangsheng, et al.
Published: (2025)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
by: Xia, Heming, et al.
Published: (2024)
by: Xia, Heming, et al.
Published: (2024)
Model Hemorrhage and the Robustness Limits of Large Language Models
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)
by: Fu, Yichao, et al.
Published: (2025)
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
by: Zhang, Libo, et al.
Published: (2024)
by: Zhang, Libo, et al.
Published: (2024)
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Label Drop for Multi-Aspect Relation Modeling in Universal Information Extraction
by: Yang, Lu, et al.
Published: (2025)
by: Yang, Lu, et al.
Published: (2025)
Speculative Decoding with a Speculative Vocabulary
by: Williams, Miles, et al.
Published: (2026)
by: Williams, Miles, et al.
Published: (2026)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
by: Zhang, Jiebin, et al.
Published: (2026)
by: Zhang, Jiebin, et al.
Published: (2026)
Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Graph-Structured Speculative Decoding
by: Gong, Zhuocheng, et al.
Published: (2024)
by: Gong, Zhuocheng, et al.
Published: (2024)
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
by: Mamou, Jonathan, et al.
Published: (2024)
by: Mamou, Jonathan, et al.
Published: (2024)
Similar Items
-
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
by: Shi, Luohe, et al.
Published: (2025) -
Reference Trustable Decoding: A Training-Free Augmentation Paradigm for Large Language Models
by: Shi, Luohe, et al.
Published: (2024) -
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
by: Tang, Zicong, et al.
Published: (2025) -
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026) -
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
by: Zhao, Yi, et al.
Published: (2025)