Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model Compression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Minjun, Choi, Jaehyeon, Yang, Hyunwoo, Kim, Jongjin, Song, Jinho, Kang, U |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning
von: Kim, Minjun, et al.
Veröffentlicht: (2026)
von: Kim, Minjun, et al.
Veröffentlicht: (2026)
Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025)
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
von: Zhou, Longsheng, et al.
Veröffentlicht: (2026)
von: Zhou, Longsheng, et al.
Veröffentlicht: (2026)
The Impact of Quantization and Pruning on Deep Reinforcement Learning Models
von: Lu, Heng, et al.
Veröffentlicht: (2024)
von: Lu, Heng, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Compression Algorithms for Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2024)
von: Park, Seungcheol, et al.
Veröffentlicht: (2024)
Locality-Aware Redundancy Pruning for LLM Depth Compression
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
von: Yun, Vincent-Daniel, et al.
Veröffentlicht: (2026)
Zero-shot Quantization: A Comprehensive Survey
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
von: Yun, Jungmin, et al.
Veröffentlicht: (2024)
Singular Value Scaling: Efficient Generative Model Compression via Pruned Weights Refinement
von: Kim, Hyeonjin, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonjin, et al.
Veröffentlicht: (2024)
DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
von: Choi, Nayoung, et al.
Veröffentlicht: (2026)
von: Choi, Nayoung, et al.
Veröffentlicht: (2026)
Model Compression using Progressive Channel Pruning
von: Guo, Jinyang, et al.
Veröffentlicht: (2025)
von: Guo, Jinyang, et al.
Veröffentlicht: (2025)
Quantization-Aware and Tensor-Compressed Training of Transformers for Natural Language Understanding
von: Yang, Zi, et al.
Veröffentlicht: (2023)
von: Yang, Zi, et al.
Veröffentlicht: (2023)
On the Compressibility of Quantized Large Language Models
von: Mao, Yu, et al.
Veröffentlicht: (2024)
von: Mao, Yu, et al.
Veröffentlicht: (2024)
LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
von: Kim, Minjun, et al.
Veröffentlicht: (2025)
Soft Label Pruning and Quantization for Large-Scale Dataset Distillation
von: Lingao, Xiao, et al.
Veröffentlicht: (2026)
von: Lingao, Xiao, et al.
Veröffentlicht: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
von: Wang, Maolin, et al.
Veröffentlicht: (2023)
Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix Compression
von: Kuroki, Kyo, et al.
Veröffentlicht: (2025)
von: Kuroki, Kyo, et al.
Veröffentlicht: (2025)
Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization
von: Monaco, Francesco Pio, et al.
Veröffentlicht: (2026)
von: Monaco, Francesco Pio, et al.
Veröffentlicht: (2026)
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
von: Janusz, Mikołaj, et al.
Veröffentlicht: (2025)
von: Janusz, Mikołaj, et al.
Veröffentlicht: (2025)
GETA-3DGS: Automatic Joint Structured Pruning and Quantization for 3D Gaussian Splatting
von: Zhang, Baobing, et al.
Veröffentlicht: (2026)
von: Zhang, Baobing, et al.
Veröffentlicht: (2026)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
Rotation Invariant Quantization for Model Compression
von: Kampeas, Joseph, et al.
Veröffentlicht: (2023)
von: Kampeas, Joseph, et al.
Veröffentlicht: (2023)
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
EPSD: Early Pruning with Self-Distillation for Efficient Model Compression
von: Chen, Dong, et al.
Veröffentlicht: (2024)
von: Chen, Dong, et al.
Veröffentlicht: (2024)
Adaptive Dataset Quantization: A New Direction for Dataset Pruning
von: Yu, Chenyue, et al.
Veröffentlicht: (2025)
von: Yu, Chenyue, et al.
Veröffentlicht: (2025)
Structured Pruning and Quantization for Learned Image Compression
von: Hossain, Md Adnan Faisal, et al.
Veröffentlicht: (2025)
von: Hossain, Md Adnan Faisal, et al.
Veröffentlicht: (2025)
HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
von: Kwon, Young D., et al.
Veröffentlicht: (2025)
von: Kwon, Young D., et al.
Veröffentlicht: (2025)
ROSAQ: Rotation-based Saliency-Aware Weight Quantization for Efficiently Compressing Large Language Models
von: Yoon, Junho, et al.
Veröffentlicht: (2025)
von: Yoon, Junho, et al.
Veröffentlicht: (2025)
DAQ: Delta-Aware Quantization for Post-Training LLM Weight Compression
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
von: Yu, Xiaoming, et al.
Veröffentlicht: (2026)
PPC-GPT: Federated Task-Specific Compression of Large Language Models via Pruning and Chain-of-Thought Distillation
von: Fan, Tao, et al.
Veröffentlicht: (2025)
von: Fan, Tao, et al.
Veröffentlicht: (2025)
Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
SwiftPrune: Hessian-Free Weight Pruning for Large Language Models
von: Kang, Yuhan, et al.
Veröffentlicht: (2025)
von: Kang, Yuhan, et al.
Veröffentlicht: (2025)
CommVQ: Commutative Vector Quantization for KV Cache Compression
von: Li, Junyan, et al.
Veröffentlicht: (2025)
von: Li, Junyan, et al.
Veröffentlicht: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
LLMs can Compress LLMs: Adaptive Pruning by Agents
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2026)
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2026)
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
von: Wagner, Moritz, et al.
Veröffentlicht: (2025)
von: Wagner, Moritz, et al.
Veröffentlicht: (2025)
Sensitivity-Guided Framework for Pruned and Quantized Reservoir Computing Accelerators
von: Jafari, Atousa, et al.
Veröffentlicht: (2026)
von: Jafari, Atousa, et al.
Veröffentlicht: (2026)
Multi-View Node Pruning for Accurate Graph Representation
von: Kim, Hanjin, et al.
Veröffentlicht: (2025)
von: Kim, Hanjin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning
von: Kim, Minjun, et al.
Veröffentlicht: (2026) -
Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression
von: Qu, Xiaoyi, et al.
Veröffentlicht: (2025) -
Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression
von: Zhou, Longsheng, et al.
Veröffentlicht: (2026) -
The Impact of Quantization and Pruning on Deep Reinforcement Learning Models
von: Lu, Heng, et al.
Veröffentlicht: (2024) -
A Comprehensive Survey of Compression Algorithms for Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2024)