HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ke, Yang, Zheng, Zhou, Zhongbin, Xue, Feng, Jiang, Zhonglin, Wang, Wenxiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
by: Kang, Yuhan, et al.
Published: (2025)
by: Kang, Yuhan, et al.
Published: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024)
by: Lu, Xudong, et al.
Published: (2024)
Orthogonal Soft Pruning for Efficient Class Unlearning
by: Gong, Qinghui, et al.
Published: (2025)
by: Gong, Qinghui, et al.
Published: (2025)
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
by: Hu, Wentao, et al.
Published: (2025)
by: Hu, Wentao, et al.
Published: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
Numerical Pruning for Efficient Autoregressive Models
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
Monte Carlo Tree Search based Space Transfer for Black-box Optimization
by: Wang, Shukuan, et al.
Published: (2024)
by: Wang, Shukuan, et al.
Published: (2024)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025)
by: Zhan, Zheng, et al.
Published: (2025)
Hessian Aware Low-Rank Perturbation for Order-Robust Continual Learning
by: Li, Jiaqi, et al.
Published: (2023)
by: Li, Jiaqi, et al.
Published: (2023)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
by: Zhao, Yushu, et al.
Published: (2025)
by: Zhao, Yushu, et al.
Published: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
by: Xie, Yanyue, et al.
Published: (2024)
by: Xie, Yanyue, et al.
Published: (2024)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
by: He, Yifei, et al.
Published: (2025)
by: He, Yifei, et al.
Published: (2025)
TopoPrune: Robust Data Pruning via Unified Latent Space Topology
by: Roy, Arjun, et al.
Published: (2026)
by: Roy, Arjun, et al.
Published: (2026)
Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model
by: Wang, Xue, et al.
Published: (2025)
by: Wang, Xue, et al.
Published: (2025)
DONOD: Efficient and Generalizable Instruction Fine-Tuning for LLMs via Model-Intrinsic Dataset Pruning
by: Hu, Jucheng, et al.
Published: (2025)
by: Hu, Jucheng, et al.
Published: (2025)
Magnitude-based Neuron Pruning for Backdoor Defens
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
Pruning-based Data Selection and Network Fusion for Efficient Deep Learning
by: Kousar, Humaira, et al.
Published: (2025)
by: Kousar, Humaira, et al.
Published: (2025)
Exploiting Adaptive Channel Pruning for Communication-Efficient Split Learning
by: Tan, Jialei, et al.
Published: (2026)
by: Tan, Jialei, et al.
Published: (2026)
Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization
by: Zhu, Haodong, et al.
Published: (2026)
by: Zhu, Haodong, et al.
Published: (2026)
Subspace-based Approximate Hessian Method for Zeroth-Order Optimization
by: Kim, Dongyoon, et al.
Published: (2025)
by: Kim, Dongyoon, et al.
Published: (2025)
End-to-End Deep Learning for Predicting Metric Space-Valued Outputs
by: Zhou, Yidong, et al.
Published: (2025)
by: Zhou, Yidong, et al.
Published: (2025)
Exploring Federated Pruning for Large Language Models
by: Guo, Pengxin, et al.
Published: (2025)
by: Guo, Pengxin, et al.
Published: (2025)
Navigating Extremes: Dynamic Sparsity in Large Output Spaces
by: Ullah, Nasib, et al.
Published: (2024)
by: Ullah, Nasib, et al.
Published: (2024)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
by: Zhang, Yike, et al.
Published: (2025)
by: Zhang, Yike, et al.
Published: (2025)
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
by: Fang, Zhiyuan, et al.
Published: (2025)
by: Fang, Zhiyuan, et al.
Published: (2025)
RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
by: Zhong, Zhengjia, et al.
Published: (2026)
by: Zhong, Zhengjia, et al.
Published: (2026)
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
by: Deng, Yonghong, et al.
Published: (2026)
by: Deng, Yonghong, et al.
Published: (2026)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
by: Kim, Junhyuck, et al.
Published: (2026)
by: Kim, Junhyuck, et al.
Published: (2026)
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces
by: Yu, Shixing, et al.
Published: (2026)
by: Yu, Shixing, et al.
Published: (2026)
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
by: Wang, Yiming, et al.
Published: (2024)
by: Wang, Yiming, et al.
Published: (2024)
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
by: Jung, Jaeheun, et al.
Published: (2025)
by: Jung, Jaeheun, et al.
Published: (2025)
Balanced Edge Pruning for Graph Anomaly Detection with Noisy Labels
by: Wang, Zhu, et al.
Published: (2024)
by: Wang, Zhu, et al.
Published: (2024)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
by: Koike-Akino, Toshiaki, et al.
Published: (2025)
by: Koike-Akino, Toshiaki, et al.
Published: (2025)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
Why Transformers Need Adam: A Hessian Perspective
by: Zhang, Yushun, et al.
Published: (2024)
by: Zhang, Yushun, et al.
Published: (2024)
Hessian-Free Online Certified Unlearning
by: Qiao, Xinbao, et al.
Published: (2024)
by: Qiao, Xinbao, et al.
Published: (2024)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
Similar Items
-
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
by: Kang, Yuhan, et al.
Published: (2025) -
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
by: Lu, Xudong, et al.
Published: (2024) -
Orthogonal Soft Pruning for Efficient Class Unlearning
by: Gong, Qinghui, et al.
Published: (2025) -
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
by: Hu, Wentao, et al.
Published: (2025) -
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
by: Yang, Cheng, et al.
Published: (2024)