BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
Fuente:
arXiv
Guardado en:
| Autores principales: | Hou, Yunlong, Zhang, Fengzhuo, Du, Cunxiao, Zhang, Xuan, Pan, Jiachun, Pang, Tianyu, Du, Chao, Tan, Vincent Y. F., Yang, Zhuoran |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
por: Yang, Penghui, et al.
Publicado: (2025)
por: Yang, Penghui, et al.
Publicado: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
por: Li, Xingyao, et al.
Publicado: (2026)
por: Li, Xingyao, et al.
Publicado: (2026)
Demystifying the Slash Pattern in Attention: The Role of RoPE
por: Cheng, Yuan, et al.
Publicado: (2026)
por: Cheng, Yuan, et al.
Publicado: (2026)
Enhancing Long Video Generation Consistency without Tuning
por: Li, Xingyao, et al.
Publicado: (2024)
por: Li, Xingyao, et al.
Publicado: (2024)
Muon Outperforms Adam in Tail-End Associative Memory Learning
por: Wang, Shuche, et al.
Publicado: (2025)
por: Wang, Shuche, et al.
Publicado: (2025)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
por: Zhang, Xuan, et al.
Publicado: (2024)
por: Zhang, Xuan, et al.
Publicado: (2024)
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
por: Xia, Heming, et al.
Publicado: (2026)
por: Xia, Heming, et al.
Publicado: (2026)
When Attention Sink Emerges in Language Models: An Empirical View
por: Gu, Xiangming, et al.
Publicado: (2024)
por: Gu, Xiangming, et al.
Publicado: (2024)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
por: Hou, Yunlong, et al.
Publicado: (2026)
por: Hou, Yunlong, et al.
Publicado: (2026)
Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
por: Hou, Yunlong, et al.
Publicado: (2024)
por: Hou, Yunlong, et al.
Publicado: (2024)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
por: Xia, Heming, et al.
Publicado: (2024)
por: Xia, Heming, et al.
Publicado: (2024)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
por: Zhang, Shuyu, et al.
Publicado: (2026)
por: Zhang, Shuyu, et al.
Publicado: (2026)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
por: Xia, Heming, et al.
Publicado: (2025)
por: Xia, Heming, et al.
Publicado: (2025)
Bandit Convex Optimization with Gradient Prediction Adaptivity
por: Wang, Shuche, et al.
Publicado: (2026)
por: Wang, Shuche, et al.
Publicado: (2026)
Indexed Minimum Empirical Divergence-Based Algorithms for Linear Bandits
por: Bian, Jie, et al.
Publicado: (2024)
por: Bian, Jie, et al.
Publicado: (2024)
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
por: Zhang, Xuan, et al.
Publicado: (2025)
por: Zhang, Xuan, et al.
Publicado: (2025)
Quantum-Enhanced Neural Contextual Bandit Algorithms
por: Huang, Yuqi, et al.
Publicado: (2026)
por: Huang, Yuqi, et al.
Publicado: (2026)
Influence Maximization via Graph Neural Bandits
por: Feng, Yuting, et al.
Publicado: (2024)
por: Feng, Yuting, et al.
Publicado: (2024)
Adversarial Combinatorial Bandits with Switching Costs
por: Dong, Yanyan, et al.
Publicado: (2024)
por: Dong, Yanyan, et al.
Publicado: (2024)
LogitSpec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation
por: Liu, Tianyu, et al.
Publicado: (2025)
por: Liu, Tianyu, et al.
Publicado: (2025)
Stochastic Bandits for Egalitarian Assignment
por: Lim, Eugene, et al.
Publicado: (2024)
por: Lim, Eugene, et al.
Publicado: (2024)
Optimal Clustering with Bandit Feedback
por: Yang, Junwen, et al.
Publicado: (2022)
por: Yang, Junwen, et al.
Publicado: (2022)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
por: Huang, Kaixuan, et al.
Publicado: (2024)
por: Huang, Kaixuan, et al.
Publicado: (2024)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
por: Wu, Hang, et al.
Publicado: (2025)
por: Wu, Hang, et al.
Publicado: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
por: Shen, Yuhao, et al.
Publicado: (2025)
por: Shen, Yuhao, et al.
Publicado: (2025)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
por: Tan, Zhendong, et al.
Publicado: (2025)
por: Tan, Zhendong, et al.
Publicado: (2025)
TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
por: Sridhar, Aditya, et al.
Publicado: (2025)
por: Sridhar, Aditya, et al.
Publicado: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
por: Liu, Hongyi, et al.
Publicado: (2025)
por: Liu, Hongyi, et al.
Publicado: (2025)
When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training
por: Wang, Haonan, et al.
Publicado: (2024)
por: Wang, Haonan, et al.
Publicado: (2024)
SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification
por: Zhang, Jiacheng, et al.
Publicado: (2026)
por: Zhang, Jiacheng, et al.
Publicado: (2026)
p-Mean Regret for Stochastic Bandits
por: Krishna, Anand, et al.
Publicado: (2024)
por: Krishna, Anand, et al.
Publicado: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
por: Xiao, Zilin, et al.
Publicado: (2024)
por: Xiao, Zilin, et al.
Publicado: (2024)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
por: Wang, Zihua, et al.
Publicado: (2025)
por: Wang, Zihua, et al.
Publicado: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
por: Shukla, Shikhar
Publicado: (2026)
por: Shukla, Shikhar
Publicado: (2026)
Efficient and Adaptive Posterior Sampling Algorithms for Bandits
por: Hu, Bingshan, et al.
Publicado: (2024)
por: Hu, Bingshan, et al.
Publicado: (2024)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
por: Ning, Zhiyuan, et al.
Publicado: (2025)
por: Ning, Zhiyuan, et al.
Publicado: (2025)
On the Optimal Regret of Locally Private Linear Contextual Bandit
por: Li, Jiachun, et al.
Publicado: (2024)
por: Li, Jiachun, et al.
Publicado: (2024)
TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification
por: Jiang, Haoyun, et al.
Publicado: (2026)
por: Jiang, Haoyun, et al.
Publicado: (2026)
SpecTr: Fast Speculative Decoding via Optimal Transport
por: Sun, Ziteng, et al.
Publicado: (2023)
por: Sun, Ziteng, et al.
Publicado: (2023)
Ejemplares similares
-
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
por: Yang, Penghui, et al.
Publicado: (2025) -
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
por: Li, Xingyao, et al.
Publicado: (2026) -
Demystifying the Slash Pattern in Attention: The Role of RoPE
por: Cheng, Yuan, et al.
Publicado: (2026) -
Enhancing Long Video Generation Consistency without Tuning
por: Li, Xingyao, et al.
Publicado: (2024) -
Muon Outperforms Adam in Tail-End Associative Memory Learning
por: Wang, Shuche, et al.
Publicado: (2025)