Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Haiduo, Yang, Fuwei, Liu, Zhenhua, Xu, Yixing, Li, Jinze, Liu, Yang, Yin, Xuanwu, Li, Dong, Ren, Pengju, Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
by: Li, Jinze, et al.
Published: (2026)
by: Li, Jinze, et al.
Published: (2026)
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
Partial Convolution Meets Visual Attention
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
by: Li, Guanchen, et al.
Published: (2025)
by: Li, Guanchen, et al.
Published: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
by: Li, Zekai, et al.
Published: (2026)
by: Li, Zekai, et al.
Published: (2026)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
by: An, Zihao, et al.
Published: (2026)
by: An, Zihao, et al.
Published: (2026)
KernelDNA: Dynamic Kernel Sharing via Decoupled Naive Adapters
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Nearly Lossless Adaptive Bit Switching
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
by: Liao, Huanxuan, et al.
Published: (2025)
by: Liao, Huanxuan, et al.
Published: (2025)
FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing
by: Li, Zekai, et al.
Published: (2024)
by: Li, Zekai, et al.
Published: (2024)
Theory-optimal Quantization Based on Flatness
by: Huang, Xiusheng, et al.
Published: (2026)
by: Huang, Xiusheng, et al.
Published: (2026)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
by: Huang, Zongle, et al.
Published: (2025)
by: Huang, Zongle, et al.
Published: (2025)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
by: McDanel, Bradley, et al.
Published: (2026)
by: McDanel, Bradley, et al.
Published: (2026)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
Partial Channel Network: Compute Fewer, Perform Better
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
by: Huang, Haiduo, et al.
Published: (2025)
by: Huang, Haiduo, et al.
Published: (2025)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
by: Li, Zeping, et al.
Published: (2024)
by: Li, Zeping, et al.
Published: (2024)
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
by: Chen, Anrui, et al.
Published: (2026)
by: Chen, Anrui, et al.
Published: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
by: Wang, Wenfeng, et al.
Published: (2025)
by: Wang, Wenfeng, et al.
Published: (2025)
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
by: Dong, Daize, et al.
Published: (2026)
by: Dong, Daize, et al.
Published: (2026)
MH-MoE: Multi-Head Mixture-of-Experts
by: Huang, Shaohan, et al.
Published: (2024)
by: Huang, Shaohan, et al.
Published: (2024)
GeGS-PCR: Effective and Robust 3D Point Cloud Registration with Two-Stage Color-Enhanced Geometric-3DGS Fusion
by: Tian, Jiayi, et al.
Published: (2026)
by: Tian, Jiayi, et al.
Published: (2026)
Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding
by: Pan, Lehan, et al.
Published: (2026)
by: Pan, Lehan, et al.
Published: (2026)
Multi-Candidate Speculative Decoding
by: Yang, Sen, et al.
Published: (2024)
by: Yang, Sen, et al.
Published: (2024)
MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios
by: Li, Shuhuai, et al.
Published: (2026)
by: Li, Shuhuai, et al.
Published: (2026)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
by: Joshi, Vinay, et al.
Published: (2025)
by: Joshi, Vinay, et al.
Published: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
by: Zhao, Boran, et al.
Published: (2025)
by: Zhao, Boran, et al.
Published: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
by: Nie, Xiaonan, et al.
Published: (2024)
by: Nie, Xiaonan, et al.
Published: (2024)
Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism
by: Cui, Chenwei, et al.
Published: (2026)
by: Cui, Chenwei, et al.
Published: (2026)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
by: Xu, Yuqi, et al.
Published: (2026)
by: Xu, Yuqi, et al.
Published: (2026)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
by: Tang, Yehui, et al.
Published: (2025)
by: Tang, Yehui, et al.
Published: (2025)
FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving
by: Liu, Qingxiu, et al.
Published: (2026)
by: Liu, Qingxiu, et al.
Published: (2026)
Similar Items
-
SpecVLM: Fast Speculative Decoding in Vision-Language Models
by: Huang, Haiduo, et al.
Published: (2025) -
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025) -
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
by: Li, Jinze, et al.
Published: (2026) -
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
by: Li, Jinze, et al.
Published: (2025) -
Partial Convolution Meets Visual Attention
by: Huang, Haiduo, et al.
Published: (2025)