Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient Implementation
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Xizi, Zhu, Jingyang, Jiang, Jingbo, Tsui, Chi-Ying |
|---|---|
| Formato: | Preprint |
| Publicado: |
2021
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Partial Knowledge Distillation for Alleviating the Inherent Inter-Class Discrepancy in Federated Learning
por: Gan, Xiaoyu, et al.
Publicado: (2024)
por: Gan, Xiaoyu, et al.
Publicado: (2024)
Accelerating Large Kernel Convolutions with Nested Winograd Transformation.pdf
por: Jiang, Jingbo, et al.
Publicado: (2021)
por: Jiang, Jingbo, et al.
Publicado: (2021)
SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
por: Su, Zeli, et al.
Publicado: (2025)
por: Su, Zeli, et al.
Publicado: (2025)
EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
por: Chen, Dong, et al.
Publicado: (2024)
por: Chen, Dong, et al.
Publicado: (2024)
Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
por: Shi, Jiang-Xin, et al.
Publicado: (2025)
por: Shi, Jiang-Xin, et al.
Publicado: (2025)
PACE: Prune-And-Compress Ensemble Models
por: Akkerman, Fabian, et al.
Publicado: (2026)
por: Akkerman, Fabian, et al.
Publicado: (2026)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
por: Yu, Tongzhou, et al.
Publicado: (2025)
por: Yu, Tongzhou, et al.
Publicado: (2025)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
por: Tang, Zichen, et al.
Publicado: (2024)
por: Tang, Zichen, et al.
Publicado: (2024)
MUC-G4: Minimal Unsat Core-Guided Incremental Verification for Deep Neural Network Compression
por: Li, Jingyang, et al.
Publicado: (2025)
por: Li, Jingyang, et al.
Publicado: (2025)
Neural Weight Compression for Language Models
por: Ryu, Jegwang, et al.
Publicado: (2025)
por: Ryu, Jegwang, et al.
Publicado: (2025)
Shapley Pruning for Neural Network Compression
por: Adamczewski, Kamil, et al.
Publicado: (2024)
por: Adamczewski, Kamil, et al.
Publicado: (2024)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
por: Zhou, Longsheng, et al.
Publicado: (2026)
por: Zhou, Longsheng, et al.
Publicado: (2026)
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
por: Kim, Kibum, et al.
Publicado: (2026)
por: Kim, Kibum, et al.
Publicado: (2026)
Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
por: Han, Xinchen, et al.
Publicado: (2026)
por: Han, Xinchen, et al.
Publicado: (2026)
Adaptive Pruning with Module Robustness Sensitivity: Balancing Compression and Robustness
por: Bai, Lincen, et al.
Publicado: (2024)
por: Bai, Lincen, et al.
Publicado: (2024)
Hyper-Compression: Model Compression via Hyperfunction
por: Fan, Fenglei, et al.
Publicado: (2024)
por: Fan, Fenglei, et al.
Publicado: (2024)
Compressing CNN models for resource-constrained systems by channel and layer pruning
por: Sadaqa, Ahmed, et al.
Publicado: (2025)
por: Sadaqa, Ahmed, et al.
Publicado: (2025)
Order of Compression: A Systematic and Optimal Sequence to Combinationally Compress CNN
por: Shen, Yingtao, et al.
Publicado: (2024)
por: Shen, Yingtao, et al.
Publicado: (2024)
Locality-Aware Redundancy Pruning for LLM Depth Compression
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
por: Yun, Vincent-Daniel, et al.
Publicado: (2026)
Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression
por: Qu, Xiaoyi, et al.
Publicado: (2025)
por: Qu, Xiaoyi, et al.
Publicado: (2025)
AutoCompress: Critical Layer Isolation for Efficient Transformer Compression
por: Thorat, Archit
Publicado: (2026)
por: Thorat, Archit
Publicado: (2026)
Spatio-Temporal Pruning for Compressed Spiking Large Language Models
por: Jiang, Yi, et al.
Publicado: (2025)
por: Jiang, Yi, et al.
Publicado: (2025)
VTrans: Accelerating Transformer Compression with Variational Information Bottleneck based Pruning
por: Dutta, Oshin, et al.
Publicado: (2024)
por: Dutta, Oshin, et al.
Publicado: (2024)
SimCert: Probabilistic Certification for Behavioral Similarity in Deep Neural Network Compression
por: Li, Jingyang, et al.
Publicado: (2026)
por: Li, Jingyang, et al.
Publicado: (2026)
Smooth Model Compression without Fine-Tuning
por: Runkel, Christina, et al.
Publicado: (2025)
por: Runkel, Christina, et al.
Publicado: (2025)
FedLAM: Low-latency Wireless Federated Learning via Layer-wise Adaptive Modulation
por: Qu, Linping, et al.
Publicado: (2025)
por: Qu, Linping, et al.
Publicado: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
por: Xu, Zhaoqi, et al.
Publicado: (2025)
por: Xu, Zhaoqi, et al.
Publicado: (2025)
Sparse Gradient Compression for Fine-Tuning Large Language Models
por: Yang, David H., et al.
Publicado: (2025)
por: Yang, David H., et al.
Publicado: (2025)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
por: Yang, Zeyu, et al.
Publicado: (2025)
por: Yang, Zeyu, et al.
Publicado: (2025)
Variance-Based Pruning for Accelerating and Compressing Trained Networks
por: Berisha, Uranik, et al.
Publicado: (2025)
por: Berisha, Uranik, et al.
Publicado: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
por: Tang, Hanlin, et al.
Publicado: (2024)
por: Tang, Hanlin, et al.
Publicado: (2024)
An Empirical Study of the Influence of Adversarial Fine-Tuning on Compressed Neural Networks
por: Thorsteinsson, Hallgrimur, et al.
Publicado: (2024)
por: Thorsteinsson, Hallgrimur, et al.
Publicado: (2024)
Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm
por: Wei, Wen-Da, et al.
Publicado: (2026)
por: Wei, Wen-Da, et al.
Publicado: (2026)
Motion-Compensated Weight Compression
por: Lamaakal, Ismail
Publicado: (2026)
por: Lamaakal, Ismail
Publicado: (2026)
Compressed Sensing: Mathematical Foundations, Implementation, and Advanced Optimization Techniques
por: Stevenson, Shane, et al.
Publicado: (2025)
por: Stevenson, Shane, et al.
Publicado: (2025)
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
por: Janusz, Mikołaj, et al.
Publicado: (2025)
por: Janusz, Mikołaj, et al.
Publicado: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
por: Xin, Jihao, et al.
Publicado: (2026)
por: Xin, Jihao, et al.
Publicado: (2026)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
por: Wagner, Moritz, et al.
Publicado: (2025)
por: Wagner, Moritz, et al.
Publicado: (2025)
EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices
por: Sanyal, Arnab, et al.
Publicado: (2025)
por: Sanyal, Arnab, et al.
Publicado: (2025)
C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
por: Bauvin, Baptiste, et al.
Publicado: (2025)
por: Bauvin, Baptiste, et al.
Publicado: (2025)
Ejemplares similares
-
Partial Knowledge Distillation for Alleviating the Inherent Inter-Class Discrepancy in Federated Learning
por: Gan, Xiaoyu, et al.
Publicado: (2024) -
Accelerating Large Kernel Convolutions with Nested Winograd Transformation.pdf
por: Jiang, Jingbo, et al.
Publicado: (2021) -
SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression
por: Su, Zeli, et al.
Publicado: (2025) -
EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
por: Chen, Dong, et al.
Publicado: (2024) -
Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
por: Shi, Jiang-Xin, et al.
Publicado: (2025)