Greedy Output Approximation: Towards Efficient Structured Pruning for LLMs Without Retraining
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jianwei, Dong, Yijun, Lei, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
by: Li, Jianwei, et al.
Published: (2023)
by: Li, Jianwei, et al.
Published: (2023)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023)
by: Zimmer, Max, et al.
Published: (2023)
Retraining as Approximate Bayesian Inference
by: Katz, Harrison
Published: (2026)
by: Katz, Harrison
Published: (2026)
Pruning Foundation Models for High Accuracy without Retraining
by: Zhao, Pu, et al.
Published: (2024)
by: Zhao, Pu, et al.
Published: (2024)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
by: Wagner, Moritz, et al.
Published: (2025)
by: Wagner, Moritz, et al.
Published: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
by: Xu, Zhaoqi, et al.
Published: (2025)
by: Xu, Zhaoqi, et al.
Published: (2025)
Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
by: Akinfaderin, Adewale, et al.
Published: (2025)
by: Akinfaderin, Adewale, et al.
Published: (2025)
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining
by: Abro, Aarash, et al.
Published: (2026)
by: Abro, Aarash, et al.
Published: (2026)
HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
All-in-One Tuning and Structural Pruning for Domain-Specific LLMs
by: Lu, Lei, et al.
Published: (2024)
by: Lu, Lei, et al.
Published: (2024)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Greedy Is Enough: Sparse Action Discovery in Agentic LLMs
by: Majumdar, Angshul
Published: (2026)
by: Majumdar, Angshul
Published: (2026)
Retrieval Augmented Anomaly Detection (RAAD): Nimble Model Adjustment Without Retraining
by: Pastoriza, Sam, et al.
Published: (2025)
by: Pastoriza, Sam, et al.
Published: (2025)
Sustainable Machine Learning Retraining: Optimizing Energy Efficiency Without Compromising Accuracy
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
by: Poenaru-Olaru, Lorena, et al.
Published: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
by: Zhang, Shuoming, et al.
Published: (2025)
by: Zhang, Shuoming, et al.
Published: (2025)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
Enhanced Structured Lasso Pruning with Class-wise Information
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Leveraging KV Similarity for Online Structured Pruning in LLMs
by: Lee, Jungmin, et al.
Published: (2025)
by: Lee, Jungmin, et al.
Published: (2025)
Towards Efficient Deep Spiking Neural Networks Construction with Spiking Activity based Pruning
by: Li, Yaxin, et al.
Published: (2024)
by: Li, Yaxin, et al.
Published: (2024)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
by: Song, Yifan, et al.
Published: (2024)
by: Song, Yifan, et al.
Published: (2024)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
by: Le, Qi, et al.
Published: (2025)
by: Le, Qi, et al.
Published: (2025)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
by: Guo, Guangfu, et al.
Published: (2025)
by: Guo, Guangfu, et al.
Published: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
by: Hasan, Adib, et al.
Published: (2024)
by: Hasan, Adib, et al.
Published: (2024)
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
by: Kovács, Ádám
Published: (2026)
by: Kovács, Ádám
Published: (2026)
Structure-Aware Automatic Channel Pruning by Searching with Graph Embedding
by: Liu, Zifan, et al.
Published: (2025)
by: Liu, Zifan, et al.
Published: (2025)
Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
by: Hyeon, Sieun, et al.
Published: (2026)
by: Hyeon, Sieun, et al.
Published: (2026)
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
Greedy Sampling Is Provably Efficient for RLHF
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
by: Yang, Jialin, et al.
Published: (2025)
by: Yang, Jialin, et al.
Published: (2025)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
PrunePath: Towards Highly Structured Sparse Language Models
by: Gu, Zhexuan, et al.
Published: (2026)
by: Gu, Zhexuan, et al.
Published: (2026)
ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
by: Xu, Xianghong, et al.
Published: (2025)
by: Xu, Xianghong, et al.
Published: (2025)
Lower Bound on the Greedy Approximation Ratio for Adaptive Submodular Cover
by: Harris, Blake, et al.
Published: (2024)
by: Harris, Blake, et al.
Published: (2024)
POP: Online Structural Pruning Enables Efficient Inference of Large Foundation Models
by: Chen, Yi, et al.
Published: (2026)
by: Chen, Yi, et al.
Published: (2026)
Towards Stable Machine Learning Model Retraining via Slowly Varying Sequences
by: Bertsimas, Dimitris, et al.
Published: (2024)
by: Bertsimas, Dimitris, et al.
Published: (2024)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
by: Li, Yanzeng, et al.
Published: (2025)
by: Li, Yanzeng, et al.
Published: (2025)
Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
Similar Items
-
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
by: Li, Jianwei, et al.
Published: (2023) -
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
by: Zimmer, Max, et al.
Published: (2023) -
Retraining as Approximate Bayesian Inference
by: Katz, Harrison
Published: (2026) -
Pruning Foundation Models for High Accuracy without Retraining
by: Zhao, Pu, et al.
Published: (2024) -
SepPrune: Structured Pruning for Efficient Deep Speech Separation
by: Li, Yuqi, et al.
Published: (2025)