Polybasic Speculative Decoding Through a Theoretical Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Ruilin, Li, Huixia, Ma, Yuexiao, Zheng, Xiawu, Chao, Fei, Xiao, Xuefeng, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AffineQuant: Affine Transformation Quantization for Large Language Models
by: Ma, Yuexiao, et al.
Published: (2024)
by: Ma, Yuexiao, et al.
Published: (2024)
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021)
by: Ma, Yuexiao, et al.
Published: (2021)
A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
by: Ma, Qingchuan, et al.
Published: (2026)
by: Ma, Qingchuan, et al.
Published: (2026)
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
A Theoretical Perspective for Speculative Decoding Algorithm
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
Benchmarking Abstract and Reasoning Abilities Through A Theoretical Perspective
by: Ma, Qingchuan, et al.
Published: (2025)
by: Ma, Qingchuan, et al.
Published: (2025)
HASTE: Training-Free Video Diffusion Acceleration via Head-Wise Adaptive Sparse Attention
by: Zheng, Xuzhe, et al.
Published: (2026)
by: Zheng, Xuzhe, et al.
Published: (2026)
Flow caching for autoregressive video generation
by: Ma, Yuexiao, et al.
Published: (2026)
by: Ma, Yuexiao, et al.
Published: (2026)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
by: Zhou, Xuwen, et al.
Published: (2026)
by: Zhou, Xuwen, et al.
Published: (2026)
A Graph Sufficiency Perspective for Neural Networks
by: Shen, Cencheng, et al.
Published: (2025)
by: Shen, Cencheng, et al.
Published: (2025)
Speculative Speculative Decoding
by: Kumar, Tanishq, et al.
Published: (2026)
by: Kumar, Tanishq, et al.
Published: (2026)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
by: Guo, Song, et al.
Published: (2024)
by: Guo, Song, et al.
Published: (2024)
Dynamic Low-Rank Sparse Adaptation for Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Decoding Speculative Decoding
by: Yan, Minghao, et al.
Published: (2024)
by: Yan, Minghao, et al.
Published: (2024)
Boosting the Cross-Architecture Generalization of Dataset Distillation through an Empirical Study
by: Zhao, Lirui, et al.
Published: (2023)
by: Zhao, Lirui, et al.
Published: (2023)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
by: Yin, Haofei, et al.
Published: (2025)
by: Yin, Haofei, et al.
Published: (2025)
Speculative Safety-Aware Decoding
by: Wang, Xuekang, et al.
Published: (2025)
by: Wang, Xuekang, et al.
Published: (2025)
Accelerating Time Series Foundation Models with Speculative Decoding
by: Subbaraman, Pranav, et al.
Published: (2025)
by: Subbaraman, Pranav, et al.
Published: (2025)
ALGOGEN: Tool-Generated Verifiable Traces for Reliable Algorithm Visualization
by: Liao, Kunpeng, et al.
Published: (2026)
by: Liao, Kunpeng, et al.
Published: (2026)
MARS: Unleashing the Power of Speculative Decoding via Margin-Aware Verification
by: Song, Jingwei, et al.
Published: (2026)
by: Song, Jingwei, et al.
Published: (2026)
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Speculative Decoding Across Languages
by: Paudel, Nirajan, et al.
Published: (2026)
by: Paudel, Nirajan, et al.
Published: (2026)
Mixture of Attentions For Speculative Decoding
by: Zimmer, Matthieu, et al.
Published: (2024)
by: Zimmer, Matthieu, et al.
Published: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification
by: Jiang, Haoyun, et al.
Published: (2026)
by: Jiang, Haoyun, et al.
Published: (2026)
Online Speculative Decoding
by: Liu, Xiaoxuan, et al.
Published: (2023)
by: Liu, Xiaoxuan, et al.
Published: (2023)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
by: Wang, Songsheng, et al.
Published: (2025)
by: Wang, Songsheng, et al.
Published: (2025)
Linear Discriminant Analysis with Gradient Optimization
by: Shen, Cencheng, et al.
Published: (2025)
by: Shen, Cencheng, et al.
Published: (2025)
High-Dimensional Independence Testing via Maximum and Average Distance Correlations
by: Shen, Cencheng, et al.
Published: (2020)
by: Shen, Cencheng, et al.
Published: (2020)
TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
by: Li, Ye, et al.
Published: (2025)
by: Li, Ye, et al.
Published: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
by: Li, Xingyao, et al.
Published: (2026)
by: Li, Xingyao, et al.
Published: (2026)
Fast Inference via Hierarchical Speculative Decoding
by: Mohri, Clara, et al.
Published: (2025)
by: Mohri, Clara, et al.
Published: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
QSpec: Speculative Decoding with Complementary Quantization Schemes
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
by: McDanel, Bradley, et al.
Published: (2026)
by: McDanel, Bradley, et al.
Published: (2026)
UniVer: A Unified Perspective for Multi-step and Multi-draft Speculative Decoding
by: Weng, Yepeng, et al.
Published: (2026)
by: Weng, Yepeng, et al.
Published: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)
by: Hou, Yunlong, et al.
Published: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
by: Fu, Yichao, et al.
Published: (2025)
by: Fu, Yichao, et al.
Published: (2025)
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Similar Items
-
AffineQuant: Affine Transformation Quantization for Large Language Models
by: Ma, Yuexiao, et al.
Published: (2024) -
Determining Layer-wise Sparsity for Large Language Models Through a Theoretical Perspective
by: Huang, Weizhong, et al.
Published: (2025) -
OMPQ: Orthogonal Mixed Precision Quantization
by: Ma, Yuexiao, et al.
Published: (2021) -
A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generation
by: Ma, Qingchuan, et al.
Published: (2026) -
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)