Adaptive Draft-Verification for Efficient Large Language Model Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Xukun, Lei, Bowen, Zhang, Ruqi, Xu, Dongkuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
di: Liu, Xukun, et al.
Pubblicazione: (2024)
di: Liu, Xukun, et al.
Pubblicazione: (2024)
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
di: Li, Jianwei, et al.
Pubblicazione: (2023)
di: Li, Jianwei, et al.
Pubblicazione: (2023)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
di: Zhang, Ziyin, et al.
Pubblicazione: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
di: Yang, Penghui, et al.
Pubblicazione: (2025)
di: Yang, Penghui, et al.
Pubblicazione: (2025)
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
di: Zhang, Situo, et al.
Pubblicazione: (2024)
di: Zhang, Situo, et al.
Pubblicazione: (2024)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
di: Sun, Chendong, et al.
Pubblicazione: (2025)
di: Sun, Chendong, et al.
Pubblicazione: (2025)
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Plato: Plan to Efficiently Decode for Large Language Model Inference
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
di: An, Zihao, et al.
Pubblicazione: (2026)
di: An, Zihao, et al.
Pubblicazione: (2026)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
di: Gao, Xiangxiang, et al.
Pubblicazione: (2024)
di: Gao, Xiangxiang, et al.
Pubblicazione: (2024)
Debiasing Large Language Models via Adaptive Causal Prompting with Sketch-of-Thought
di: Li, Bowen, et al.
Pubblicazione: (2026)
di: Li, Bowen, et al.
Pubblicazione: (2026)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
di: Lin, Chun-Hsien, et al.
Pubblicazione: (2024)
di: Lin, Chun-Hsien, et al.
Pubblicazione: (2024)
Mixture of Decoding: An Attention-Inspired Adaptive Decoding Strategy to Mitigate Hallucinations in Large Vision-Language Models
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
di: Chen, Xinlong, et al.
Pubblicazione: (2025)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
di: Cao, Jiaqi, et al.
Pubblicazione: (2025)
di: Cao, Jiaqi, et al.
Pubblicazione: (2025)
APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation
di: Zheng, Tianyu, et al.
Pubblicazione: (2026)
di: Zheng, Tianyu, et al.
Pubblicazione: (2026)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
di: Sun, Ryan, et al.
Pubblicazione: (2024)
di: Sun, Ryan, et al.
Pubblicazione: (2024)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
di: Wu, Zhaoxuan, et al.
Pubblicazione: (2025)
di: Wu, Zhaoxuan, et al.
Pubblicazione: (2025)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
di: Yang, Ze, et al.
Pubblicazione: (2024)
di: Yang, Ze, et al.
Pubblicazione: (2024)
Semantic-guided Diverse Decoding for Large Language Model
di: Shi, Weijie, et al.
Pubblicazione: (2025)
di: Shi, Weijie, et al.
Pubblicazione: (2025)
Efficient Attention Mechanisms for Large Language Models: A Survey
di: Sun, Yutao, et al.
Pubblicazione: (2025)
di: Sun, Yutao, et al.
Pubblicazione: (2025)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
di: Hu, Shijing, et al.
Pubblicazione: (2025)
di: Hu, Shijing, et al.
Pubblicazione: (2025)
Variation in Verification: Understanding Verification Dynamics in Large Language Models
di: Zhou, Yefan, et al.
Pubblicazione: (2025)
di: Zhou, Yefan, et al.
Pubblicazione: (2025)
Exploring and Improving Drafts in Blockwise Parallel Decoding
di: Kim, Taehyeon, et al.
Pubblicazione: (2024)
di: Kim, Taehyeon, et al.
Pubblicazione: (2024)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
di: Liu, Zheyuan, et al.
Pubblicazione: (2025)
di: Liu, Zheyuan, et al.
Pubblicazione: (2025)
Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models
di: Liu, Peijie, et al.
Pubblicazione: (2025)
di: Liu, Peijie, et al.
Pubblicazione: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
di: Hong, Fenglu, et al.
Pubblicazione: (2025)
di: Hong, Fenglu, et al.
Pubblicazione: (2025)
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit
di: Wang, Kangyu, et al.
Pubblicazione: (2025)
di: Wang, Kangyu, et al.
Pubblicazione: (2025)
Collaborative Stance Detection via Small-Large Language Model Consistency Verification
di: Yan, Yu, et al.
Pubblicazione: (2025)
di: Yan, Yu, et al.
Pubblicazione: (2025)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
di: Reddy, Avinash, et al.
Pubblicazione: (2026)
di: Reddy, Avinash, et al.
Pubblicazione: (2026)
Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs
di: Thompson, Travis, et al.
Pubblicazione: (2025)
di: Thompson, Travis, et al.
Pubblicazione: (2025)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
di: Goel, Raghavv, et al.
Pubblicazione: (2024)
ARS: Adaptive Reasoning Suppression for Efficient Large Reasoning Language Models
di: Zheng, Dongqi
Pubblicazione: (2025)
di: Zheng, Dongqi
Pubblicazione: (2025)
One Brain, Omni Modalities: Towards Unified Non-Invasive Brain Decoding with Large Language Models
di: Tang, Changli, et al.
Pubblicazione: (2026)
di: Tang, Changli, et al.
Pubblicazione: (2026)
Generation Meets Verification: Accelerating Large Language Model Inference with Smart Parallel Auto-Correct Decoding
di: Yi, Hanling, et al.
Pubblicazione: (2024)
di: Yi, Hanling, et al.
Pubblicazione: (2024)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
di: Su, Xin, et al.
Pubblicazione: (2026)
di: Su, Xin, et al.
Pubblicazione: (2026)
LLM-Barber: Block-Aware Rebuilder for Sparsity Mask in One-Shot for Large Language Models
di: Su, Yupeng, et al.
Pubblicazione: (2024)
di: Su, Yupeng, et al.
Pubblicazione: (2024)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
di: Hong, Ruixin, et al.
Pubblicazione: (2023)
di: Hong, Ruixin, et al.
Pubblicazione: (2023)
Efficient Large Language Models: A Survey
di: Wan, Zhongwei, et al.
Pubblicazione: (2023)
di: Wan, Zhongwei, et al.
Pubblicazione: (2023)
Documenti analoghi
-
ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph
di: Liu, Xukun, et al.
Pubblicazione: (2024) -
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
di: Li, Jianwei, et al.
Pubblicazione: (2023) -
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
di: Zhang, Ziyin, et al.
Pubblicazione: (2024) -
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
di: Yang, Penghui, et al.
Pubblicazione: (2025) -
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
di: Zhang, Situo, et al.
Pubblicazione: (2024)