SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Zeli, Zhang, Ziyin, Zhang, Wenzheng, Liu, Zhou, Xu, Guixian, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
by: Su, Zeli, et al.
Published: (2026)
by: Su, Zeli, et al.
Published: (2026)
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
by: Su, Zeli, et al.
Published: (2025)
by: Su, Zeli, et al.
Published: (2025)
FedPaI: Achieving Extreme Sparsity in Federated Learning via Pruning at Initialization
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
by: Xu, Guixian, et al.
Published: (2025)
by: Xu, Guixian, et al.
Published: (2025)
Efficient LLMs with AMP: Attention Heads and MLP Pruning
by: Mugnaini, Leandro Giusti, et al.
Published: (2025)
by: Mugnaini, Leandro Giusti, et al.
Published: (2025)
pQuant: Towards Effective Low-Bit Language Models via Decoupled Linear Quantization-Aware Training
by: Zhang, Wenzheng, et al.
Published: (2026)
by: Zhang, Wenzheng, et al.
Published: (2026)
Symmetry Induces Structure and Constraint of Learning
by: Ziyin, Liu
Published: (2023)
by: Ziyin, Liu
Published: (2023)
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
by: Wang, Hanrui, et al.
Published: (2020)
by: Wang, Hanrui, et al.
Published: (2020)
EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
by: Chen, Dong, et al.
Published: (2024)
by: Chen, Dong, et al.
Published: (2024)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
by: Xu, Guixian, et al.
Published: (2026)
by: Xu, Guixian, et al.
Published: (2026)
Three Mechanisms of Feature Learning in a Linear Network
by: Xu, Yizhou, et al.
Published: (2024)
by: Xu, Yizhou, et al.
Published: (2024)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
by: Zhou, Longsheng, et al.
Published: (2026)
by: Zhou, Longsheng, et al.
Published: (2026)
ReNCE: Learning to Reason by Noise Contrastive Estimation
by: Zhang, Wenzheng, et al.
Published: (2026)
by: Zhang, Wenzheng, et al.
Published: (2026)
Community-Centric Graph Unlearning
by: Li, Yi, et al.
Published: (2024)
by: Li, Yi, et al.
Published: (2024)
Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient Implementation
by: Chen, Xizi, et al.
Published: (2021)
by: Chen, Xizi, et al.
Published: (2021)
Learning Multiplex Representations on Text-Attributed Graphs with One Language Model Encoder
by: Jin, Bowen, et al.
Published: (2023)
by: Jin, Bowen, et al.
Published: (2023)
Exploring Learning Complexity for Efficient Downstream Dataset Pruning
by: Jiang, Wenyu, et al.
Published: (2024)
by: Jiang, Wenyu, et al.
Published: (2024)
Divide, Specialize, and Route: A New Approach to Efficient Ensemble Learning
by: Piwko, Jakub, et al.
Published: (2025)
by: Piwko, Jakub, et al.
Published: (2025)
LongFlow: Efficient KV Cache Compression for Reasoning Models
by: Su, Yi, et al.
Published: (2026)
by: Su, Yi, et al.
Published: (2026)
Multi-Head Spectral-Adaptive Graph Anomaly Detection
by: Cao, Qingyue, et al.
Published: (2025)
by: Cao, Qingyue, et al.
Published: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
by: Xu, Zhaoqi, et al.
Published: (2025)
by: Xu, Zhaoqi, et al.
Published: (2025)
FL-PLAS: Federated Learning with Partial Layer Aggregation for Backdoor Defense Against High-Ratio Malicious Clients
by: Zhang, Jianyi, et al.
Published: (2025)
by: Zhang, Jianyi, et al.
Published: (2025)
Differential Informed Auto-Encoder
by: Zhang, Jinrui
Published: (2024)
by: Zhang, Jinrui
Published: (2024)
Numerical Pruning for Efficient Autoregressive Models
by: Shen, Xuan, et al.
Published: (2024)
by: Shen, Xuan, et al.
Published: (2024)
Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient Transformers
by: Gabetni, Firas, et al.
Published: (2025)
by: Gabetni, Firas, et al.
Published: (2025)
Toward Fair Graph Neural Networks Via Dual-Teacher Knowledge Distillation
by: Li, Chengyu, et al.
Published: (2024)
by: Li, Chengyu, et al.
Published: (2024)
An Equivariance Toolbox for Learning Dynamics
by: Yang, Yongyi, et al.
Published: (2025)
by: Yang, Yongyi, et al.
Published: (2025)
A New Convergence Analysis of Plug-and-Play Proximal Gradient Descent Under Prior Mismatch
by: Xu, Guixian, et al.
Published: (2026)
by: Xu, Guixian, et al.
Published: (2026)
Learning to Route Among Specialized Experts for Zero-Shot Generalization
by: Muqeeth, Mohammed, et al.
Published: (2024)
by: Muqeeth, Mohammed, et al.
Published: (2024)
Single-Stage Huffman Encoder for ML Compression
by: Agrawal, Aditya, et al.
Published: (2026)
by: Agrawal, Aditya, et al.
Published: (2026)
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
by: Xu, Jingjing, et al.
Published: (2024)
by: Xu, Jingjing, et al.
Published: (2024)
EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
by: Xing, Xingrun, et al.
Published: (2025)
by: Xing, Xingrun, et al.
Published: (2025)
CipherPrune: Efficient and Scalable Private Transformer Inference
by: Zhang, Yancheng, et al.
Published: (2025)
by: Zhang, Yancheng, et al.
Published: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
Towards Efficient Automatic Self-Pruning of Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
by: Li, Pingzhi, et al.
Published: (2023)
by: Li, Pingzhi, et al.
Published: (2023)
Remove Symmetries to Control Model Expressivity and Improve Optimization
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
LieTrunc-QNN: Lie Algebra Truncation and Quantum Expressivity Phase Transition from LiePrune to Provably Stable Quantum Neural Networks
by: Shao, Haijian, et al.
Published: (2026)
by: Shao, Haijian, et al.
Published: (2026)
Adaptive Graph Auto-Encoder for General Data Clustering
by: Li, Xuelong, et al.
Published: (2020)
by: Li, Xuelong, et al.
Published: (2020)
HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
Similar Items
-
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
by: Su, Zeli, et al.
Published: (2026) -
Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages
by: Su, Zeli, et al.
Published: (2025) -
FedPaI: Achieving Extreme Sparsity in Federated Learning via Pruning at Initialization
by: Wang, Haonan, et al.
Published: (2025) -
CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China
by: Xu, Guixian, et al.
Published: (2025) -
Efficient LLMs with AMP: Attention Heads and MLP Pruning
by: Mugnaini, Leandro Giusti, et al.
Published: (2025)