TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Sibo, Fu, Jinyuan, Xie, Zhongle, Shou, Lidan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
Acceptance Dynamics Across Cognitive Domains in Speculative Decoding
von: Mahmoud, Saif
Veröffentlicht: (2026)
von: Mahmoud, Saif
Veröffentlicht: (2026)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
von: Fang, Min, et al.
Veröffentlicht: (2025)
von: Fang, Min, et al.
Veröffentlicht: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
von: Huo, Feiye, et al.
Veröffentlicht: (2025)
von: Huo, Feiye, et al.
Veröffentlicht: (2025)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
von: Shou, Yitong, et al.
Veröffentlicht: (2026)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
RASD: Retrieval-Augmented Speculative Decoding
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
PAD: Personalized Alignment of LLMs at Decoding-Time
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2024)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
Alignment-Enhanced Decoding:Defending via Token-Level Adaptive Refining of Probability Distributions
von: Liu, Quan, et al.
Veröffentlicht: (2024)
von: Liu, Quan, et al.
Veröffentlicht: (2024)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
Entropy-UID: A Method for Optimizing Information Density
von: Shou, Xinpeng
Veröffentlicht: (2025)
von: Shou, Xinpeng
Veröffentlicht: (2025)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025) -
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025) -
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025) -
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024) -
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)