On Speculative Decoding for Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gagrani, Mukul, Goel, Raghavv, Jeon, Wonseok, Park, Junyoung, Lee, Mingu, Lott, Christopher |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
ConFu: Contemplate the Future for Better Speculative Sampling
von: Qin, Zongyue, et al.
Veröffentlicht: (2026)
von: Qin, Zongyue, et al.
Veröffentlicht: (2026)
CAOTE: KV Cache Selection for LLMs via Attention Output Error-Based Token Eviction
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
von: Gautam, Aayush, et al.
Veröffentlicht: (2026)
von: Gautam, Aayush, et al.
Veröffentlicht: (2026)
AdaEDL: Early Draft Stopping for Speculative Decoding of Large Language Models via an Entropy-based Lower Bound on Token Acceptance Probability
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2024)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2024)
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments
von: Park, Junyoung, et al.
Veröffentlicht: (2025)
von: Park, Junyoung, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
von: Jones, Dalton, et al.
Veröffentlicht: (2026)
von: Jones, Dalton, et al.
Veröffentlicht: (2026)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
REST: Retrieval-Based Speculative Decoding
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
Benchmarking the Energy Savings with Speculative Decoding Strategies
von: Dutta, Rohit, et al.
Veröffentlicht: (2026)
von: Dutta, Rohit, et al.
Veröffentlicht: (2026)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Self-Speculative Biased Decoding for Faster Re-Translation
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
von: Goel, Naman
Veröffentlicht: (2023)
von: Goel, Naman
Veröffentlicht: (2023)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
von: Park, Jean, et al.
Veröffentlicht: (2024)
von: Park, Jean, et al.
Veröffentlicht: (2024)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
Pruning as a Defense: Reducing Memorization in Large Language Models
von: Gupta, Mansi, et al.
Veröffentlicht: (2025)
von: Gupta, Mansi, et al.
Veröffentlicht: (2025)
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
von: Zhao, Weilin, et al.
Veröffentlicht: (2025)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2024) -
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024) -
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025) -
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2025) -
ConFu: Contemplate the Future for Better Speculative Sampling
von: Qin, Zongyue, et al.
Veröffentlicht: (2026)