Progressive Binarization with Semi-Structured Pruning for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Xianglong, Zhang, Tianao, Li, Zhiteng, Qin, Haotong, Zhang, Yulun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025)
by: Zhang, Tianao, et al.
Published: (2025)
ARB-LLM: Alternating Refined Binarizations for Large Language Models
by: Li, Zhiteng, et al.
Published: (2024)
by: Li, Zhiteng, et al.
Published: (2024)
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
by: Yan, Xianglong, et al.
Published: (2025)
by: Yan, Xianglong, et al.
Published: (2025)
PT$^2$-LLM: Post-Training Ternarization for Large Language Models
by: Yan, Xianglong, et al.
Published: (2025)
by: Yan, Xianglong, et al.
Published: (2025)
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)
by: Yan, Xianglong, et al.
Published: (2026)
SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
by: Bao, Chengzhu, et al.
Published: (2026)
by: Bao, Chengzhu, et al.
Published: (2026)
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
by: Chen, Hong, et al.
Published: (2024)
by: Chen, Hong, et al.
Published: (2024)
Low-bit Model Quantization for Deep Neural Networks: A Survey
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
by: Huang, Wei, et al.
Published: (2023)
by: Huang, Wei, et al.
Published: (2023)
BinaryHPE: 3D Human Pose and Shape Estimation via Binarization
by: Li, Zhiteng, et al.
Published: (2023)
by: Li, Zhiteng, et al.
Published: (2023)
FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
BiVM: Accurate Binarized Neural Network for Efficient Video Matting
by: Qin, Haotong, et al.
Published: (2025)
by: Qin, Haotong, et al.
Published: (2025)
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
by: Qin, Haotong, et al.
Published: (2024)
by: Qin, Haotong, et al.
Published: (2024)
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
An empirical study of LLaMA3 quantization: from LLMs to MLLMs
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
BiDense: Binarization for Dense Prediction
by: Yin, Rui, et al.
Published: (2024)
by: Yin, Rui, et al.
Published: (2024)
Binarized Simplicial Convolutional Neural Networks
by: Yan, Yi, et al.
Published: (2024)
by: Yan, Yi, et al.
Published: (2024)
Binarized Diffusion Model for Image Super-Resolution
by: Chen, Zheng, et al.
Published: (2024)
by: Chen, Zheng, et al.
Published: (2024)
Beyond Manually Designed Pruning Policies with Second-Level Performance Prediction: A Pruning Framework for LLMs
by: Ma, Zuxin, et al.
Published: (2025)
by: Ma, Zuxin, et al.
Published: (2025)
AI Progress Should Be Measured by Capability-Per-Resource, Not Scale Alone: A Framework for Gradient-Guided Resource Allocation in LLMs
by: McCoy, David, et al.
Published: (2025)
by: McCoy, David, et al.
Published: (2025)
Graph Construction with Flexible Nodes for Traffic Demand Prediction
by: Hou, Jinyan, et al.
Published: (2024)
by: Hou, Jinyan, et al.
Published: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
by: Li, Guanchen, et al.
Published: (2025)
by: Li, Guanchen, et al.
Published: (2025)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
by: Liu, Lawrence, et al.
Published: (2025)
by: Liu, Lawrence, et al.
Published: (2025)
Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models with Structured Pruning in Resource-Limited Clients
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
by: Ye, Zijian, et al.
Published: (2025)
by: Ye, Zijian, et al.
Published: (2025)
Two-Stage Regularization-Based Structured Pruning for LLMs
by: Feng, Mingkuan, et al.
Published: (2025)
by: Feng, Mingkuan, et al.
Published: (2025)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
2DQuant: Low-bit Post-Training Quantization for Image Super-Resolution
by: Liu, Kai, et al.
Published: (2024)
by: Liu, Kai, et al.
Published: (2024)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
BiPFT: Binary Pre-trained Foundation Transformer with Low-rank Estimation of Binarization Residual Polynomials
by: Xing, Xingrun, et al.
Published: (2023)
by: Xing, Xingrun, et al.
Published: (2023)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
by: Fu, Yao, et al.
Published: (2025)
by: Fu, Yao, et al.
Published: (2025)
An Empirical Study of Qwen3 Quantization
by: Zheng, Xingyu, et al.
Published: (2025)
by: Zheng, Xingyu, et al.
Published: (2025)
PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs
by: Yu, Tongzhou, et al.
Published: (2025)
by: Yu, Tongzhou, et al.
Published: (2025)
Understanding Neural Network Binarization with Forward and Backward Proximal Quantizers
by: Lu, Yiwei, et al.
Published: (2024)
by: Lu, Yiwei, et al.
Published: (2024)
Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning
by: Luo, Xinjian, et al.
Published: (2020)
by: Luo, Xinjian, et al.
Published: (2020)
Deterministic Differentiable Structured Pruning for Large Language Models
by: Huang, Weiyu, et al.
Published: (2026)
by: Huang, Weiyu, et al.
Published: (2026)
On Pruning State-Space LLMs
by: Ghattas, Tamer, et al.
Published: (2025)
by: Ghattas, Tamer, et al.
Published: (2025)
Similar Items
-
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
by: Zhang, Tianao, et al.
Published: (2025) -
ARB-LLM: Alternating Refined Binarizations for Large Language Models
by: Li, Zhiteng, et al.
Published: (2024) -
ReCalKV: Low-Rank KV Cache Compression via Head Reordering and Offline Calibration
by: Yan, Xianglong, et al.
Published: (2025) -
PT$^2$-LLM: Post-Training Ternarization for Large Language Models
by: Yan, Xianglong, et al.
Published: (2025) -
D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs
by: Yan, Xianglong, et al.
Published: (2026)