Speculative Safety-Aware Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xuekang, Zhu, Shengyu, Cheng, Xueqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning
von: Su, Tongrui, et al.
Veröffentlicht: (2025)
von: Su, Tongrui, et al.
Veröffentlicht: (2025)
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
von: He, Liang, et al.
Veröffentlicht: (2026)
von: He, Liang, et al.
Veröffentlicht: (2026)
QSpec: Speculative Decoding with Complementary Quantization Schemes
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
von: Xie, Zhinan, et al.
Veröffentlicht: (2025)
von: Xie, Zhinan, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
von: Li, Ye, et al.
Veröffentlicht: (2025)
von: Li, Ye, et al.
Veröffentlicht: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
STree: Speculative Tree Decoding for Hybrid State-Space Models
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
When Drafts Evolve: Speculative Decoding Meets Online Learning
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
Self-Speculative Biased Decoding for Faster Re-Translation
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Benchmarking the Energy Savings with Speculative Decoding Strategies
von: Dutta, Rohit, et al.
Veröffentlicht: (2026)
von: Dutta, Rohit, et al.
Veröffentlicht: (2026)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
REST: Retrieval-Based Speculative Decoding
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
von: Wu, Tianyu, et al.
Veröffentlicht: (2026)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
von: Cha, Seongjin, et al.
Veröffentlicht: (2026)
von: Cha, Seongjin, et al.
Veröffentlicht: (2026)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning
von: Su, Tongrui, et al.
Veröffentlicht: (2025) -
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
von: He, Liang, et al.
Veröffentlicht: (2026) -
QSpec: Speculative Decoding with Complementary Quantization Schemes
von: Zhao, Juntao, et al.
Veröffentlicht: (2024) -
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
von: Park, Jihoon, et al.
Veröffentlicht: (2025) -
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)