EAGLE-Pangu: Accelerator-Safe Tree Speculative Decoding on Ascend NPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Han, Chang, Hu, Yijie, Liu, Jingling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Learned Knowledge in LoRA Adapters Through Efficient Contrastive Decoding on Ascend NPUs
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025)
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025)
Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
von: Guan, Yue, et al.
Veröffentlicht: (2025)
von: Guan, Yue, et al.
Veröffentlicht: (2025)
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
von: Yin, Yichun, et al.
Veröffentlicht: (2025)
von: Yin, Yichun, et al.
Veröffentlicht: (2025)
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
von: Taghian, Mehran, et al.
Veröffentlicht: (2026)
von: Taghian, Mehran, et al.
Veröffentlicht: (2026)
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
von: Tang, Yehui, et al.
Veröffentlicht: (2025)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2026)
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Block Verification Accelerates Speculative Decoding
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
Speculative Decoding Across Languages
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
Dato: A Task-Based Programming Model for Dataflow Accelerators
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
von: Fang, Shihan, et al.
Veröffentlicht: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
Out-of-Vocabulary Sampling Boosts Speculative Decoding
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
UniVer: A Unified Perspective for Multi-step and Multi-draft Speculative Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2026)
von: Weng, Yepeng, et al.
Veröffentlicht: (2026)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Accelerating PayPal's Commerce Agent with Speculative Decoding: An Empirical Study on EAGLE3 with Fine-Tuned Nemotron Models
von: Qin, Ally, et al.
Veröffentlicht: (2026)
von: Qin, Ally, et al.
Veröffentlicht: (2026)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
von: Li, Yuhui, et al.
Veröffentlicht: (2024)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
Optimized Speculative Sampling for GPU Hardware Accelerators
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
Speculative Decoding for Verilog: Speed and Quality, All in One
von: Xu, Changran, et al.
Veröffentlicht: (2025)
von: Xu, Changran, et al.
Veröffentlicht: (2025)
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
von: Samarin, Alexander, et al.
Veröffentlicht: (2026)
von: Samarin, Alexander, et al.
Veröffentlicht: (2026)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2025)
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2025)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024)
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing Learned Knowledge in LoRA Adapters Through Efficient Contrastive Decoding on Ascend NPUs
von: Heisler, Morgan Lindsay, et al.
Veröffentlicht: (2025) -
Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
von: Guan, Yue, et al.
Veröffentlicht: (2025) -
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs
von: Yin, Yichun, et al.
Veröffentlicht: (2025) -
HiFloat4 Format for Language Model Pre-training on Ascend NPUs
von: Taghian, Mehran, et al.
Veröffentlicht: (2026) -
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs
von: Tang, Yehui, et al.
Veröffentlicht: (2025)