MSQ: Memory-Efficient Bit Sparsification Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Seokho, Yoon, Seoyeon, Kim, Jinhee, Wang, Dongwei, Jeon, Kang Eun, Yang, Huanrui, Ko, Jong Hwan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026)
by: Wang, Dongwei, et al.
Published: (2026)
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025)
by: Kim, Jinhee, et al.
Published: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)
by: Kang, Do Yeong, et al.
Published: (2025)
Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization
by: Wang, Dongwei, et al.
Published: (2024)
by: Wang, Dongwei, et al.
Published: (2024)
BitSnap: Checkpoint Sparsification and Quantization in LLM Training
by: Peng, Yanxin, et al.
Published: (2025)
by: Peng, Yanxin, et al.
Published: (2025)
Row-Column Hybrid Grouping for Fault-Resilient Multi-Bit Weight Representation on IMC Arrays
by: Jeon, Kang Eun, et al.
Published: (2025)
by: Jeon, Kang Eun, et al.
Published: (2025)
Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification
by: Huang, Hong, et al.
Published: (2026)
by: Huang, Hong, et al.
Published: (2026)
One-Bit Quantization and Sparsification for Multiclass Linear Classification with Strong Regularization
by: Ghane, Reza, et al.
Published: (2024)
by: Ghane, Reza, et al.
Published: (2024)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
by: Deng, Jianing, et al.
Published: (2026)
by: Deng, Jianing, et al.
Published: (2026)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
by: Chong, Hyochan, et al.
Published: (2026)
by: Chong, Hyochan, et al.
Published: (2026)
Efficient Asynchronous Federated Learning with Sparsification and Quantization
by: Jia, Juncheng, et al.
Published: (2023)
by: Jia, Juncheng, et al.
Published: (2023)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
by: Kim, Taesoo, et al.
Published: (2025)
by: Kim, Taesoo, et al.
Published: (2025)
Efficient Distributed Training through Gradient Compression with Sparsification and Quantization Techniques
by: Singh, Shruti, et al.
Published: (2024)
by: Singh, Shruti, et al.
Published: (2024)
DyBit: Dynamic Bit-Precision Numbers for Efficient Quantized Neural Network Inference
by: Zhou, Jiajun, et al.
Published: (2023)
by: Zhou, Jiajun, et al.
Published: (2023)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
by: Park, Jungwoo, et al.
Published: (2025)
by: Park, Jungwoo, et al.
Published: (2025)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
by: Lee, Banseok, et al.
Published: (2025)
by: Lee, Banseok, et al.
Published: (2025)
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
by: Bang, Wonjun, et al.
Published: (2025)
by: Bang, Wonjun, et al.
Published: (2025)
BAQ: Efficient Bit Allocation Quantization for Large Language Models
by: Zhang, Chao, et al.
Published: (2025)
by: Zhang, Chao, et al.
Published: (2025)
PaAno: Patch-Based Representation Learning for Time-Series Anomaly Detection
by: Park, Jinju, et al.
Published: (2026)
by: Park, Jinju, et al.
Published: (2026)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
by: Cho, Yoonjun, et al.
Published: (2026)
by: Cho, Yoonjun, et al.
Published: (2026)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
Where and How to Enhance: Discovering Bit-Width Contribution for Mixed Precision Quantization
by: Kang, Haidong, et al.
Published: (2025)
by: Kang, Haidong, et al.
Published: (2025)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
by: Choi, Kanghyun, et al.
Published: (2024)
by: Choi, Kanghyun, et al.
Published: (2024)
MOMEMTO: Patch-based Memory Gate Model in Time Series Foundation Model
by: Yoon, Samuel, et al.
Published: (2025)
by: Yoon, Samuel, et al.
Published: (2025)
Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators
by: Zhou, Wenyong, et al.
Published: (2025)
by: Zhou, Wenyong, et al.
Published: (2025)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
by: Jeon, Hyesung, et al.
Published: (2024)
by: Jeon, Hyesung, et al.
Published: (2024)
RDIS: Random Drop Imputation with Self-Training for Incomplete Time Series Data
by: Choi, Tae-Min, et al.
Published: (2020)
by: Choi, Tae-Min, et al.
Published: (2020)
Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization
by: Zhao, Jiayu, et al.
Published: (2026)
by: Zhao, Jiayu, et al.
Published: (2026)
Efficient Unbiased Sparsification
by: Barnes, Leighton, et al.
Published: (2024)
by: Barnes, Leighton, et al.
Published: (2024)
LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning
by: Guo, Han, et al.
Published: (2023)
by: Guo, Han, et al.
Published: (2023)
Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
by: Kim, Jaemin, et al.
Published: (2026)
by: Kim, Jaemin, et al.
Published: (2026)
SenDaL: An Effective and Efficient Calibration Framework of Low-Cost Sensors for Daily Life
by: Ahn, Seokho, et al.
Published: (2025)
by: Ahn, Seokho, et al.
Published: (2025)
One-Bit Quantization for Random Features Models
by: Akhtiamov, Danil, et al.
Published: (2025)
by: Akhtiamov, Danil, et al.
Published: (2025)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
by: Park, Yeonhong, et al.
Published: (2024)
by: Park, Yeonhong, et al.
Published: (2024)
A Carbon Tracking Model for Federated Learning: Impact of Quantization and Sparsification
by: Barbieri, Luca, et al.
Published: (2023)
by: Barbieri, Luca, et al.
Published: (2023)
Similar Items
-
TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit Precision
by: Kim, Jinhee, et al.
Published: (2025) -
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
by: Wang, Dongwei, et al.
Published: (2026) -
Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset Sampling
by: Kim, Jinhee, et al.
Published: (2025) -
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025) -
MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing Architectures
by: Kang, Do Yeong, et al.
Published: (2025)