Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jinze, Xu, Yixing, Huang, Haiduo, Yin, Xuanwu, Li, Dong, Ngai, Edith C. H., Barsoum, Emad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2026)
von: Li, Jinze, et al.
Veröffentlicht: (2026)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)
von: An, Zihao, et al.
Veröffentlicht: (2026)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
MSWA: Refining Local Attention with Multi-ScaleWindow Attention
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
von: Xu, Yixing, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
von: Li, Zekai, et al.
Veröffentlicht: (2026)
von: Li, Zekai, et al.
Veröffentlicht: (2026)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing
von: Li, Zekai, et al.
Veröffentlicht: (2024)
von: Li, Zekai, et al.
Veröffentlicht: (2024)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
von: Li, Guanchen, et al.
Veröffentlicht: (2025)
von: Li, Guanchen, et al.
Veröffentlicht: (2025)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
von: Li, Zeping, et al.
Veröffentlicht: (2024)
von: Li, Zeping, et al.
Veröffentlicht: (2024)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
Theory-optimal Quantization Based on Flatness
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
von: Yang, Shuo, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs
von: Xiao, Sibo, et al.
Veröffentlicht: (2025)
von: Xiao, Sibo, et al.
Veröffentlicht: (2025)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
von: Ray, Pretam, et al.
Veröffentlicht: (2026)
Partial Convolution Meets Visual Attention
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
von: Li, Jinze, et al.
Veröffentlicht: (2025) -
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2026) -
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)