Earley-Driven Dynamic Pruning for Efficient Structured Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Xintong, Wei, Chi, Tian, Minghao, Ni, Shiwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
di: Zhang, Yike, et al.
Pubblicazione: (2025)
di: Zhang, Yike, et al.
Pubblicazione: (2025)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
di: Qin, Jiayu, et al.
Pubblicazione: (2025)
di: Qin, Jiayu, et al.
Pubblicazione: (2025)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
di: Yu, Fengming, et al.
Pubblicazione: (2025)
di: Yu, Fengming, et al.
Pubblicazione: (2025)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
di: Pan, Rui, et al.
Pubblicazione: (2025)
di: Pan, Rui, et al.
Pubblicazione: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
di: Fu, Qichen, et al.
Pubblicazione: (2024)
di: Fu, Qichen, et al.
Pubblicazione: (2024)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
di: Simonds, Toby
Pubblicazione: (2025)
di: Simonds, Toby
Pubblicazione: (2025)
Adaptive Pruning for Large Language Models with Structural Importance Awareness
di: Zheng, Haotian, et al.
Pubblicazione: (2024)
di: Zheng, Haotian, et al.
Pubblicazione: (2024)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
di: Le, Qi, et al.
Pubblicazione: (2025)
di: Le, Qi, et al.
Pubblicazione: (2025)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
di: Wang, Ziyan, et al.
Pubblicazione: (2025)
di: Wang, Ziyan, et al.
Pubblicazione: (2025)
Scaling Inference-Efficient Language Models
di: Bian, Song, et al.
Pubblicazione: (2025)
di: Bian, Song, et al.
Pubblicazione: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
di: Vincenti, Jort, et al.
Pubblicazione: (2024)
di: Vincenti, Jort, et al.
Pubblicazione: (2024)
Two-Stage Regularization-Based Structured Pruning for LLMs
di: Feng, Mingkuan, et al.
Pubblicazione: (2025)
di: Feng, Mingkuan, et al.
Pubblicazione: (2025)
Text Quality-Based Pruning for Efficient Training of Language Models
di: Sharma, Vasu, et al.
Pubblicazione: (2024)
di: Sharma, Vasu, et al.
Pubblicazione: (2024)
Sample-aware Adaptive Structured Pruning for Large Language Models
di: Kong, Jun, et al.
Pubblicazione: (2025)
di: Kong, Jun, et al.
Pubblicazione: (2025)
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding
di: Xiao, Zhongyu, et al.
Pubblicazione: (2026)
di: Xiao, Zhongyu, et al.
Pubblicazione: (2026)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
di: Zhu, Hourun, et al.
Pubblicazione: (2025)
di: Zhu, Hourun, et al.
Pubblicazione: (2025)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
di: Zhou, Longsheng, et al.
Pubblicazione: (2026)
di: Zhou, Longsheng, et al.
Pubblicazione: (2026)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
di: Sandri, Fabrizio, et al.
Pubblicazione: (2025)
di: Sandri, Fabrizio, et al.
Pubblicazione: (2025)
Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment
di: Liu, Jun, et al.
Pubblicazione: (2024)
di: Liu, Jun, et al.
Pubblicazione: (2024)
BESA: Pruning Large Language Models with Blockwise Parameter-Efficient Sparsity Allocation
di: Xu, Peng, et al.
Pubblicazione: (2024)
di: Xu, Peng, et al.
Pubblicazione: (2024)
From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models
di: Wang, Ziyan, et al.
Pubblicazione: (2025)
di: Wang, Ziyan, et al.
Pubblicazione: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
di: Patel, Dev, et al.
Pubblicazione: (2025)
di: Patel, Dev, et al.
Pubblicazione: (2025)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
di: Pei, Zehua, et al.
Pubblicazione: (2024)
di: Pei, Zehua, et al.
Pubblicazione: (2024)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
di: Gautam, Aayush, et al.
Pubblicazione: (2025)
di: Gautam, Aayush, et al.
Pubblicazione: (2025)
A Simple and Effective Pruning Approach for Large Language Models
di: Sun, Mingjie, et al.
Pubblicazione: (2023)
di: Sun, Mingjie, et al.
Pubblicazione: (2023)
RelayLLM: Efficient Reasoning via Collaborative Decoding
di: Huang, Chengsong, et al.
Pubblicazione: (2026)
di: Huang, Chengsong, et al.
Pubblicazione: (2026)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
di: Reddy, Avinash, et al.
Pubblicazione: (2026)
di: Reddy, Avinash, et al.
Pubblicazione: (2026)
An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
di: Xia, Mengzhou, et al.
Pubblicazione: (2023)
di: Xia, Mengzhou, et al.
Pubblicazione: (2023)
Pruning Foundation Models for High Accuracy without Retraining
di: Zhao, Pu, et al.
Pubblicazione: (2024)
di: Zhao, Pu, et al.
Pubblicazione: (2024)
IDEA Prune: An Integrated Enlarge-and-Prune Pipeline in Generative Language Model Pretraining
di: Li, Yixiao, et al.
Pubblicazione: (2025)
di: Li, Yixiao, et al.
Pubblicazione: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
di: Hu, Yuezhou, et al.
Pubblicazione: (2025)
di: Hu, Yuezhou, et al.
Pubblicazione: (2025)
The Structure of Relation Decoding Linear Operators in Large Language Models
di: Christ, Miranda Anna, et al.
Pubblicazione: (2025)
di: Christ, Miranda Anna, et al.
Pubblicazione: (2025)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
di: Geng, Saibo, et al.
Pubblicazione: (2023)
di: Geng, Saibo, et al.
Pubblicazione: (2023)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
di: Lu, Xudong, et al.
Pubblicazione: (2024)
di: Lu, Xudong, et al.
Pubblicazione: (2024)
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models
di: Das, Rocktim Jyoti, et al.
Pubblicazione: (2023)
di: Das, Rocktim Jyoti, et al.
Pubblicazione: (2023)
On the Limitations of Language Targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning
di: Kurz, Simon, et al.
Pubblicazione: (2024)
di: Kurz, Simon, et al.
Pubblicazione: (2024)
Traversal Verification for Speculative Tree Decoding
di: Weng, Yepeng, et al.
Pubblicazione: (2025)
di: Weng, Yepeng, et al.
Pubblicazione: (2025)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
di: Hu, Haiquan, et al.
Pubblicazione: (2025)
di: Hu, Haiquan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
di: Dong, Harry, et al.
Pubblicazione: (2024) -
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
di: Zhang, Yike, et al.
Pubblicazione: (2025) -
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
di: Qin, Jiayu, et al.
Pubblicazione: (2025) -
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
di: Yu, Fengming, et al.
Pubblicazione: (2025) -
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
di: Pan, Rui, et al.
Pubblicazione: (2025)