GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Selim, Suh, Il hong, Kim, Yeseong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
di: Kim, Junhan, et al.
Pubblicazione: (2026)
di: Kim, Junhan, et al.
Pubblicazione: (2026)
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
di: Bhattacharyya, Chaitali, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Chaitali, et al.
Pubblicazione: (2025)
Low-Rank Quantization-Aware Training for LLMs
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
di: Zhou, Sifan, et al.
Pubblicazione: (2025)
DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
di: Mikaelyan, Liana, et al.
Pubblicazione: (2025)
di: Mikaelyan, Liana, et al.
Pubblicazione: (2025)
Continuous Approximations for Improving Quantization Aware Training of LLMs
di: Li, He, et al.
Pubblicazione: (2024)
di: Li, He, et al.
Pubblicazione: (2024)
Calibration and Transformation-Free Weight-Only LLMs Quantization via Dynamic Grouping
di: Zheng, Xinzhe, et al.
Pubblicazione: (2025)
di: Zheng, Xinzhe, et al.
Pubblicazione: (2025)
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
di: Deng, Yanxia, et al.
Pubblicazione: (2025)
di: Deng, Yanxia, et al.
Pubblicazione: (2025)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
LoopQ: Quantization for Recursive Transformers
di: Fang, Rui, et al.
Pubblicazione: (2026)
di: Fang, Rui, et al.
Pubblicazione: (2026)
Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs
di: Qiao, Ye, et al.
Pubblicazione: (2025)
di: Qiao, Ye, et al.
Pubblicazione: (2025)
Exploiting Boosting in Hyperdimensional Computing for Enhanced Reliability in Healthcare
di: Jeong, SungHeon, et al.
Pubblicazione: (2024)
di: Jeong, SungHeon, et al.
Pubblicazione: (2024)
CoreQ: Learning-Free Mismatch Correction and Successive Rounding for Quantization
di: Cha, Seohyeon, et al.
Pubblicazione: (2026)
di: Cha, Seohyeon, et al.
Pubblicazione: (2026)
D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
di: Li, Junlin, et al.
Pubblicazione: (2026)
di: Li, Junlin, et al.
Pubblicazione: (2026)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
di: Lv, Mengtao, et al.
Pubblicazione: (2025)
di: Lv, Mengtao, et al.
Pubblicazione: (2025)
Universal Approximation Theorem of Deep Q-Networks
di: Qi, Qian
Pubblicazione: (2025)
di: Qi, Qian
Pubblicazione: (2025)
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
di: Chen, Kejia, et al.
Pubblicazione: (2025)
di: Chen, Kejia, et al.
Pubblicazione: (2025)
DeltaDQ: Ultra-High Delta Compression for Fine-Tuned LLMs via Group-wise Dropout and Separate Quantization
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
di: Jiang, Yanfeng, et al.
Pubblicazione: (2024)
AMiD: Knowledge Distillation for LLMs with $α$-mixture Assistant Distribution
di: Shin, Donghyeok, et al.
Pubblicazione: (2025)
di: Shin, Donghyeok, et al.
Pubblicazione: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
di: Gao, Yuting, et al.
Pubblicazione: (2025)
di: Gao, Yuting, et al.
Pubblicazione: (2025)
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
di: Du, Ally Yalei, et al.
Pubblicazione: (2024)
di: Du, Ally Yalei, et al.
Pubblicazione: (2024)
Q-realign: Piggybacking Realignment on Quantization for Safe and Efficient LLM Deployment
di: Tan, Qitao, et al.
Pubblicazione: (2026)
di: Tan, Qitao, et al.
Pubblicazione: (2026)
A Shared Low-Rank Adaptation Approach to Personalized RLHF
di: Liu, Renpu, et al.
Pubblicazione: (2025)
di: Liu, Renpu, et al.
Pubblicazione: (2025)
LinkQ: An LLM-Assisted Visual Interface for Knowledge Graph Question-Answering
di: Li, Harry, et al.
Pubblicazione: (2024)
di: Li, Harry, et al.
Pubblicazione: (2024)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
di: Wu, Zhuguanyu, et al.
Pubblicazione: (2025)
di: Wu, Zhuguanyu, et al.
Pubblicazione: (2025)
GroupRank: A Groupwise Paradigm for Effective and Efficient Passage Reranking with LLMs
di: Long, Meixiu, et al.
Pubblicazione: (2025)
di: Long, Meixiu, et al.
Pubblicazione: (2025)
Matrix Low-Rank Approximation For Policy Gradient Methods
di: Rozada, Sergio, et al.
Pubblicazione: (2024)
di: Rozada, Sergio, et al.
Pubblicazione: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2024)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
di: Ye, Rongguang, et al.
Pubblicazione: (2025)
QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs
di: Mishra, Himanshu, et al.
Pubblicazione: (2026)
di: Mishra, Himanshu, et al.
Pubblicazione: (2026)
Navigation with QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning
di: Canesse, Alexi, et al.
Pubblicazione: (2024)
di: Canesse, Alexi, et al.
Pubblicazione: (2024)
State Rank Dynamics in Linear Attention LLMs
di: Sun, Ao, et al.
Pubblicazione: (2026)
di: Sun, Ao, et al.
Pubblicazione: (2026)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
di: Lee, Geonho, et al.
Pubblicazione: (2024)
di: Lee, Geonho, et al.
Pubblicazione: (2024)
Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement Learning
di: Park, Jongchan, et al.
Pubblicazione: (2025)
di: Park, Jongchan, et al.
Pubblicazione: (2025)
HeadQ: Model-Visible Distortion and Score-Space Correction for KV-Cache Quantization
di: Williams, Jorge L. Ruiz
Pubblicazione: (2026)
di: Williams, Jorge L. Ruiz
Pubblicazione: (2026)
Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
di: Yang, Yuchen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026) -
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
di: Kim, Junhan, et al.
Pubblicazione: (2026) -
FineScope : SAE-guided Data Selection Enables Domain Specific LLM Pruning and Finetuning
di: Bhattacharyya, Chaitali, et al.
Pubblicazione: (2025) -
Low-Rank Quantization-Aware Training for LLMs
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024) -
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
di: Zhou, Sifan, et al.
Pubblicazione: (2025)