Greedy Output Approximation: Towards Efficient Structured Pruning for LLMs Without Retraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jianwei, Dong, Yijun, Lei, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
von: Zimmer, Max, et al.
Veröffentlicht: (2023)
von: Zimmer, Max, et al.
Veröffentlicht: (2023)
Retraining as Approximate Bayesian Inference
von: Katz, Harrison
Veröffentlicht: (2026)
von: Katz, Harrison
Veröffentlicht: (2026)
Pruning Foundation Models for High Accuracy without Retraining
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
von: Zhao, Pu, et al.
Veröffentlicht: (2024)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
von: Wagner, Moritz, et al.
Veröffentlicht: (2025)
von: Wagner, Moritz, et al.
Veröffentlicht: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
von: Xu, Zhaoqi, et al.
Veröffentlicht: (2025)
von: Xu, Zhaoqi, et al.
Veröffentlicht: (2025)
Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
Spectral Unforgetting: Post-Hoc Recovery of Damaged Capabilities Without Retraining
von: Abro, Aarash, et al.
Veröffentlicht: (2026)
von: Abro, Aarash, et al.
Veröffentlicht: (2026)
HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
von: Li, Ke, et al.
Veröffentlicht: (2025)
von: Li, Ke, et al.
Veröffentlicht: (2025)
All-in-One Tuning and Structural Pruning for Domain-Specific LLMs
von: Lu, Lei, et al.
Veröffentlicht: (2024)
von: Lu, Lei, et al.
Veröffentlicht: (2024)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
von: Li, Meng, et al.
Veröffentlicht: (2026)
von: Li, Meng, et al.
Veröffentlicht: (2026)
Greedy Is Enough: Sparse Action Discovery in Agentic LLMs
von: Majumdar, Angshul
Veröffentlicht: (2026)
von: Majumdar, Angshul
Veröffentlicht: (2026)
Retrieval Augmented Anomaly Detection (RAAD): Nimble Model Adjustment Without Retraining
von: Pastoriza, Sam, et al.
Veröffentlicht: (2025)
von: Pastoriza, Sam, et al.
Veröffentlicht: (2025)
Sustainable Machine Learning Retraining: Optimizing Energy Efficiency Without Compromising Accuracy
von: Poenaru-Olaru, Lorena, et al.
Veröffentlicht: (2025)
von: Poenaru-Olaru, Lorena, et al.
Veröffentlicht: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
von: Yu, Tongzhou, et al.
Veröffentlicht: (2025)
von: Yu, Tongzhou, et al.
Veröffentlicht: (2025)
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
von: Zhang, Shuoming, et al.
Veröffentlicht: (2025)
von: Zhang, Shuoming, et al.
Veröffentlicht: (2025)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
Enhanced Structured Lasso Pruning with Class-wise Information
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
von: Liu, Xiang, et al.
Veröffentlicht: (2025)
Leveraging KV Similarity for Online Structured Pruning in LLMs
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
von: Lee, Jungmin, et al.
Veröffentlicht: (2025)
Towards Efficient Deep Spiking Neural Networks Construction with Spiking Activity based Pruning
von: Li, Yaxin, et al.
Veröffentlicht: (2024)
von: Li, Yaxin, et al.
Veröffentlicht: (2024)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
von: Le, Qi, et al.
Veröffentlicht: (2025)
von: Le, Qi, et al.
Veröffentlicht: (2025)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
von: Ao, Shuang, et al.
Veröffentlicht: (2025)
von: Ao, Shuang, et al.
Veröffentlicht: (2025)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
von: Guo, Guangfu, et al.
Veröffentlicht: (2025)
von: Guo, Guangfu, et al.
Veröffentlicht: (2025)
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
von: Hasan, Adib, et al.
Veröffentlicht: (2024)
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
von: Kovács, Ádám
Veröffentlicht: (2026)
von: Kovács, Ádám
Veröffentlicht: (2026)
Structure-Aware Automatic Channel Pruning by Searching with Graph Embedding
von: Liu, Zifan, et al.
Veröffentlicht: (2025)
von: Liu, Zifan, et al.
Veröffentlicht: (2025)
Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
von: Hyeon, Sieun, et al.
Veröffentlicht: (2026)
Two-Stage Regularization-Based Structured Pruning for LLMs
von: Feng, Mingkuan, et al.
Veröffentlicht: (2025)
von: Feng, Mingkuan, et al.
Veröffentlicht: (2025)
Greedy Sampling Is Provably Efficient for RLHF
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
von: Yang, Jialin, et al.
Veröffentlicht: (2025)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
von: Li, Yan, et al.
Veröffentlicht: (2025)
von: Li, Yan, et al.
Veröffentlicht: (2025)
PrunePath: Towards Highly Structured Sparse Language Models
von: Gu, Zhexuan, et al.
Veröffentlicht: (2026)
von: Gu, Zhexuan, et al.
Veröffentlicht: (2026)
ZeroCard: Cardinality Estimation with Zero Dependence on Target Databases -- No Data, No Query, No Retraining
von: Xu, Xianghong, et al.
Veröffentlicht: (2025)
von: Xu, Xianghong, et al.
Veröffentlicht: (2025)
Lower Bound on the Greedy Approximation Ratio for Adaptive Submodular Cover
von: Harris, Blake, et al.
Veröffentlicht: (2024)
von: Harris, Blake, et al.
Veröffentlicht: (2024)
POP: Online Structural Pruning Enables Efficient Inference of Large Foundation Models
von: Chen, Yi, et al.
Veröffentlicht: (2026)
von: Chen, Yi, et al.
Veröffentlicht: (2026)
Towards Stable Machine Learning Model Retraining via Slowly Varying Sequences
von: Bertsimas, Dimitris, et al.
Veröffentlicht: (2024)
von: Bertsimas, Dimitris, et al.
Veröffentlicht: (2024)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
von: Li, Yanzeng, et al.
Veröffentlicht: (2025)
von: Li, Yanzeng, et al.
Veröffentlicht: (2025)
Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
von: Wang, Zijun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
von: Li, Jianwei, et al.
Veröffentlicht: (2023) -
PERP: Rethinking the Prune-Retrain Paradigm in the Era of LLMs
von: Zimmer, Max, et al.
Veröffentlicht: (2023) -
Retraining as Approximate Bayesian Inference
von: Katz, Harrison
Veröffentlicht: (2026) -
Pruning Foundation Models for High Accuracy without Retraining
von: Zhao, Pu, et al.
Veröffentlicht: (2024) -
SepPrune: Structured Pruning for Efficient Deep Speech Separation
von: Li, Yuqi, et al.
Veröffentlicht: (2025)