Multi-Drafter Speculative Decoding with Alignment Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Taehyeon, Jung, Hojung, Yun, Se-Young |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
von: Kim, Taehyeon, et al.
Veröffentlicht: (2023)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2023)
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024)
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024)
Coupling without Communication and Drafter-Invariant Speculative Decoding
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
Guiding Reasoning in Small Language Models with LLM Assistance
von: Kim, Yujin, et al.
Veröffentlicht: (2025)
von: Kim, Yujin, et al.
Veröffentlicht: (2025)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
von: Ramakrishnan, Ramchalam Kinattinkara, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Ramchalam Kinattinkara, et al.
Veröffentlicht: (2025)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
von: Lee, Gihun, et al.
Veröffentlicht: (2024)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
Revisiting Early-Learning Regularization When Federated Learning Meets Noisy Labels
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2024)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
Multi-Candidate Speculative Decoding
von: Yang, Sen, et al.
Veröffentlicht: (2024)
von: Yang, Sen, et al.
Veröffentlicht: (2024)
3-Model Speculative Decoding
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
von: Kim, Sungkyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungkyun, et al.
Veröffentlicht: (2025)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
von: Koh, Woosung, et al.
Veröffentlicht: (2024)
Improving Multi-candidate Speculative Decoding
von: Lu, Xiaofan, et al.
Veröffentlicht: (2024)
von: Lu, Xiaofan, et al.
Veröffentlicht: (2024)
Bayesian Multi-Task Transfer Learning for Soft Prompt Tuning
von: Lee, Haeju, et al.
Veröffentlicht: (2024)
von: Lee, Haeju, et al.
Veröffentlicht: (2024)
Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models
von: Park, Youngrok, et al.
Veröffentlicht: (2025)
von: Park, Youngrok, et al.
Veröffentlicht: (2025)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Speculative Contrastive Decoding
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
A Multi-Model Adaptation of Speculative Decoding for Classification
von: Roy, Somnath, et al.
Veröffentlicht: (2025)
von: Roy, Somnath, et al.
Veröffentlicht: (2025)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
von: Ho, Namgyu, et al.
Veröffentlicht: (2024)
von: Ho, Namgyu, et al.
Veröffentlicht: (2024)
READER: Retrieval-Assisted Drafter for Efficient LLM Inference
von: Divilkovskiy, Maxim, et al.
Veröffentlicht: (2025)
von: Divilkovskiy, Maxim, et al.
Veröffentlicht: (2025)
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
von: Lee, Segyu, et al.
Veröffentlicht: (2026)
von: Lee, Segyu, et al.
Veröffentlicht: (2026)
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
von: Oh, Jihwan, et al.
Veröffentlicht: (2026)
von: Oh, Jihwan, et al.
Veröffentlicht: (2026)
Graph-Structured Speculative Decoding
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
von: Lv, Kai, et al.
Veröffentlicht: (2025)
von: Lv, Kai, et al.
Veröffentlicht: (2025)
Steering Pretrained Drafters during Speculative Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
Towards Optimal Multi-draft Speculative Decoding
von: Hu, Zhengmian, et al.
Veröffentlicht: (2025)
von: Hu, Zhengmian, et al.
Veröffentlicht: (2025)
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024) -
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025) -
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
von: Kim, Taehyeon, et al.
Veröffentlicht: (2023) -
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025) -
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)