SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Park, Yeonsik, Kim, Hyeonseong, Choi, Seungkyu |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring
par: Lee, Dongyoung, et autres
Publié: (2025)
par: Lee, Dongyoung, et autres
Publié: (2025)
LQER: Low-Rank Quantization Error Reconstruction for LLMs
par: Zhang, Cheng, et autres
Publié: (2024)
par: Zhang, Cheng, et autres
Publié: (2024)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
par: Cho, Yoonjun, et autres
Publié: (2026)
par: Cho, Yoonjun, et autres
Publié: (2026)
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
par: Kim, Hee-Sung, et autres
Publié: (2026)
par: Kim, Hee-Sung, et autres
Publié: (2026)
Low-Rank Quantization-Aware Training for LLMs
par: Bondarenko, Yelysei, et autres
Publié: (2024)
par: Bondarenko, Yelysei, et autres
Publié: (2024)
Saliency-Aware Regularized Quantization Calibration for Large Language Models
par: Zhao, Yanlong, et autres
Publié: (2026)
par: Zhao, Yanlong, et autres
Publié: (2026)
Beta-Sigma VAE: Separating beta and decoder variance in Gaussian variational autoencoder
par: Kim, Seunghwan, et autres
Publié: (2024)
par: Kim, Seunghwan, et autres
Publié: (2024)
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
par: Gul, Hongyaoxing, et autres
Publié: (2026)
par: Gul, Hongyaoxing, et autres
Publié: (2026)
DecDEC: A Systems Approach to Advancing Low-Bit LLM Quantization
par: Park, Yeonhong, et autres
Publié: (2024)
par: Park, Yeonhong, et autres
Publié: (2024)
Saliency-Aware Model Merging
par: Park, Jungin, et autres
Publié: (2026)
par: Park, Jungin, et autres
Publié: (2026)
Low-Rank Correction for Quantized LLMs
par: Scetbon, Meyer, et autres
Publié: (2024)
par: Scetbon, Meyer, et autres
Publié: (2024)
QERA: an Analytical Framework for Quantization Error Reconstruction
par: Zhang, Cheng, et autres
Publié: (2024)
par: Zhang, Cheng, et autres
Publié: (2024)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
par: Lee, Geonho, et autres
Publié: (2024)
par: Lee, Geonho, et autres
Publié: (2024)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
par: Huang, Wei, et autres
Publié: (2024)
par: Huang, Wei, et autres
Publié: (2024)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
par: Li, Yuhang, et autres
Publié: (2024)
par: Li, Yuhang, et autres
Publié: (2024)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
par: Tang, Pingzhi, et autres
Publié: (2026)
par: Tang, Pingzhi, et autres
Publié: (2026)
Rethinking Residual Errors in Compensation-based LLM Quantization
par: Li, Shuaiting, et autres
Publié: (2026)
par: Li, Shuaiting, et autres
Publié: (2026)
Saliency Assisted Quantization for Neural Networks
par: Rezabeyk, Elmira Mousa, et autres
Publié: (2024)
par: Rezabeyk, Elmira Mousa, et autres
Publié: (2024)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
par: Lee, Seoungsub, et autres
Publié: (2026)
par: Lee, Seoungsub, et autres
Publié: (2026)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
par: Abitante, João Vitor Boer, et autres
Publié: (2026)
par: Abitante, João Vitor Boer, et autres
Publié: (2026)
UltraSketchLLM: Saliency-Driven Sketching for Ultra-Low Bit LLM Compression
par: Zou, Sunan, et autres
Publié: (2025)
par: Zou, Sunan, et autres
Publié: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
par: Kim, Minsu, et autres
Publié: (2025)
par: Kim, Minsu, et autres
Publié: (2025)
Learning to Continually Learn with the Bayesian Principle
par: Lee, Soochan, et autres
Publié: (2024)
par: Lee, Soochan, et autres
Publié: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
par: Loeschcke, Sebastian, et autres
Publié: (2024)
par: Loeschcke, Sebastian, et autres
Publié: (2024)
Saliency-Aware Regularized Graph Neural Network
par: Pei, Wenjie, et autres
Publié: (2024)
par: Pei, Wenjie, et autres
Publié: (2024)
Toward Architecture-Agnostic Local Control of Posterior Collapse in VAEs
par: Song, Hyunsoo, et autres
Publié: (2025)
par: Song, Hyunsoo, et autres
Publié: (2025)
HLQ: Fast and Efficient Backpropagation via Hadamard Low-rank Quantization
par: Kim, Seonggon, et autres
Publié: (2024)
par: Kim, Seonggon, et autres
Publié: (2024)
FAQ: Mitigating Quantization Error via Regenerating Calibration Data with Family-Aware Quantization
par: Xiao, Haiyang, et autres
Publié: (2026)
par: Xiao, Haiyang, et autres
Publié: (2026)
FinLoRA: Finetuning Quantized Financial Large Language Models Using Low-Rank Adaptation
par: Wang, Dannong, et autres
Publié: (2024)
par: Wang, Dannong, et autres
Publié: (2024)
Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control
par: Park, Seongmin, et autres
Publié: (2024)
par: Park, Seongmin, et autres
Publié: (2024)
ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
par: Zhao, Weibo, et autres
Publié: (2024)
par: Zhao, Weibo, et autres
Publié: (2024)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
par: Bouquet, Yann, et autres
Publié: (2026)
par: Bouquet, Yann, et autres
Publié: (2026)
Stabilizing Native Low-Rank LLM Pretraining
par: Janson, Paul, et autres
Publié: (2026)
par: Janson, Paul, et autres
Publié: (2026)
SPaRSe-TIME: Saliency-Projected Low-Rank Temporal Modeling for Efficient and Interpretable Time Series Prediction
par: Shahriar, K. A.
Publié: (2026)
par: Shahriar, K. A.
Publié: (2026)
LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
par: Gu, Hongyaoxing, et autres
Publié: (2026)
par: Gu, Hongyaoxing, et autres
Publié: (2026)
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
par: Gu, Hongyaoxing, et autres
Publié: (2026)
par: Gu, Hongyaoxing, et autres
Publié: (2026)
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
par: Huang, Beichen, et autres
Publié: (2025)
par: Huang, Beichen, et autres
Publié: (2025)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
par: Cho, Yoonjun, et autres
Publié: (2025)
par: Cho, Yoonjun, et autres
Publié: (2025)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
par: Cai, Wen-Pu, et autres
Publié: (2024)
par: Cai, Wen-Pu, et autres
Publié: (2024)
ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression
par: Yu, Wneya, et autres
Publié: (2026)
par: Yu, Wneya, et autres
Publié: (2026)
Documents similaires
-
Qrazor: Reliable and Effortless 4-bit LLM Quantization by Significant Data Razoring
par: Lee, Dongyoung, et autres
Publié: (2025) -
LQER: Low-Rank Quantization Error Reconstruction for LLMs
par: Zhang, Cheng, et autres
Publié: (2024) -
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
par: Cho, Yoonjun, et autres
Publié: (2026) -
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
par: Kim, Hee-Sung, et autres
Publié: (2026) -
Low-Rank Quantization-Aware Training for LLMs
par: Bondarenko, Yelysei, et autres
Publié: (2024)