ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Shuzhang, Yang, Zebin, Li, Meng, Gong, Ruihao, Wang, Runsheng, Huang, Ru |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2024)
by: Zhong, Shuzhang, et al.
Published: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2025)
by: Zhong, Shuzhang, et al.
Published: (2025)
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
by: Zhong, Shuzhang, et al.
Published: (2026)
by: Zhong, Shuzhang, et al.
Published: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
by: Jo, Hyun-rae, et al.
Published: (2024)
by: Jo, Hyun-rae, et al.
Published: (2024)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
by: Ding, Xueying, et al.
Published: (2025)
by: Ding, Xueying, et al.
Published: (2025)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
by: Wei, Linye, et al.
Published: (2025)
by: Wei, Linye, et al.
Published: (2025)
Earley-Driven Dynamic Pruning for Efficient Structured Decoding
by: Sun, Xintong, et al.
Published: (2025)
by: Sun, Xintong, et al.
Published: (2025)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
by: Qin, Zongyue, et al.
Published: (2024)
by: Qin, Zongyue, et al.
Published: (2024)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
by: Taniguchi, Rei, et al.
Published: (2026)
by: Taniguchi, Rei, et al.
Published: (2026)
Attention Sink Forges Native MoE in Attention Layers: Sink-Aware Training to Address Head Collapse
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Parallel Token Prediction for Language Models
by: Draxler, Felix, et al.
Published: (2025)
by: Draxler, Felix, et al.
Published: (2025)
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers
by: Yang, Zebin, et al.
Published: (2024)
by: Yang, Zebin, et al.
Published: (2024)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
by: Xiao, Zhongyu, et al.
Published: (2026)
by: Xiao, Zhongyu, et al.
Published: (2026)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
by: Guan, Yue, et al.
Published: (2025)
by: Guan, Yue, et al.
Published: (2025)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
by: Rodionov, Gleb, et al.
Published: (2025)
by: Rodionov, Gleb, et al.
Published: (2025)
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning
by: Bi, Jiaxi, et al.
Published: (2026)
by: Bi, Jiaxi, et al.
Published: (2026)
ProCut: LLM Prompt Compression via Attribution Estimation
by: Xu, Zhentao, et al.
Published: (2025)
by: Xu, Zhentao, et al.
Published: (2025)
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
by: Guo, Zhenyuan, et al.
Published: (2026)
by: Guo, Zhenyuan, et al.
Published: (2026)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
by: Li, Jinze, et al.
Published: (2025)
by: Li, Jinze, et al.
Published: (2025)
RelayLLM: Efficient Reasoning via Collaborative Decoding
by: Huang, Chengsong, et al.
Published: (2026)
by: Huang, Chengsong, et al.
Published: (2026)
Improving Diffusion Language Model Decoding through Joint Search in Generation Order and Token Space
by: Shen, Yangyi, et al.
Published: (2026)
by: Shen, Yangyi, et al.
Published: (2026)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
by: Zhang, Yike, et al.
Published: (2025)
by: Zhang, Yike, et al.
Published: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
by: Oba, Daisuke, et al.
Published: (2026)
by: Oba, Daisuke, et al.
Published: (2026)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
by: Ziashahabi, Amir, et al.
Published: (2025)
by: Ziashahabi, Amir, et al.
Published: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
by: Yu, Fengming, et al.
Published: (2025)
by: Yu, Fengming, et al.
Published: (2025)
Towards Auto-Regressive Next-Token Prediction: In-Context Learning Emerges from Generalization
by: Gong, Zixuan, et al.
Published: (2025)
by: Gong, Zixuan, et al.
Published: (2025)
Efficient Vision-Language Reasoning via Adaptive Token Pruning
by: Li, Xue, et al.
Published: (2025)
by: Li, Xue, et al.
Published: (2025)
Exploring and Improving Drafts in Blockwise Parallel Decoding
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
Similar Items
-
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2024) -
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
by: Zhong, Shuzhang, et al.
Published: (2025) -
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration
by: Zhong, Shuzhang, et al.
Published: (2026) -
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024) -
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
by: Jo, Hyun-rae, et al.
Published: (2024)