LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
Fuente:
arXiv
Salvato in:
| Autori principali: | Gu, Hongyaoxing, Hu, Lijuan, Yu, Liye, Li, Haowei, Liu, Fangfang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026)
A method of using RSVD in residual calculation of LowBit GEMM
di: Gu, Hongyaoxing
Pubblicazione: (2024)
di: Gu, Hongyaoxing
Pubblicazione: (2024)
PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
di: Wang, Sheng, et al.
Pubblicazione: (2024)
di: Wang, Sheng, et al.
Pubblicazione: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)
LoRAP: Low-Rank Aggregation Prompting for Quantized Graph Neural Networks Training
di: Liu, Chenyu, et al.
Pubblicazione: (2026)
di: Liu, Chenyu, et al.
Pubblicazione: (2026)
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations
di: Blumenberg, Patrick, et al.
Pubblicazione: (2025)
di: Blumenberg, Patrick, et al.
Pubblicazione: (2025)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
di: Bouquet, Yann, et al.
Pubblicazione: (2026)
di: Bouquet, Yann, et al.
Pubblicazione: (2026)
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
di: Huang, Beichen, et al.
Pubblicazione: (2025)
di: Huang, Beichen, et al.
Pubblicazione: (2025)
LRAMM -- Low precision approximates GEMM via RSVD
di: Gu, Hongyaoxing
Pubblicazione: (2024)
di: Gu, Hongyaoxing
Pubblicazione: (2024)
FinLoRA: Finetuning Quantized Financial Large Language Models Using Low-Rank Adaptation
di: Wang, Dannong, et al.
Pubblicazione: (2024)
di: Wang, Dannong, et al.
Pubblicazione: (2024)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
Pushing the Limits of Block Rotations in Post-Training Quantization
di: Sanjeet, Sai, et al.
Pubblicazione: (2026)
di: Sanjeet, Sai, et al.
Pubblicazione: (2026)
Low-Rank Correction for Quantized LLMs
di: Scetbon, Meyer, et al.
Pubblicazione: (2024)
di: Scetbon, Meyer, et al.
Pubblicazione: (2024)
Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart
di: Yu, Chengting, et al.
Pubblicazione: (2024)
di: Yu, Chengting, et al.
Pubblicazione: (2024)
LoRA-GA: Low-Rank Adaptation with Gradient Approximation
di: Wang, Shaowen, et al.
Pubblicazione: (2024)
di: Wang, Shaowen, et al.
Pubblicazione: (2024)
SmoothRot: Combining Channel-Wise Scaling and Rotation for Quantization-Friendly LLMs
di: Czakó, Patrik, et al.
Pubblicazione: (2025)
di: Czakó, Patrik, et al.
Pubblicazione: (2025)
Block Rotation is All You Need for MXFP4 Quantization
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
BoRA: Towards More Expressive Low-Rank Adaptation with Block Diversity
di: Li, Shiwei, et al.
Pubblicazione: (2025)
di: Li, Shiwei, et al.
Pubblicazione: (2025)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
BlockEcho: Retaining Long-Range Dependencies for Imputing Block-Wise Missing Data
di: Han, Qiao, et al.
Pubblicazione: (2024)
di: Han, Qiao, et al.
Pubblicazione: (2024)
LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
di: Mirzaei, Amir Reza, et al.
Pubblicazione: (2025)
di: Mirzaei, Amir Reza, et al.
Pubblicazione: (2025)
Merging LoRAs like Playing LEGO: Pushing the Modularity of LoRA to Extremes Through Rank-Wise Clustering
di: Zhao, Ziyu, et al.
Pubblicazione: (2024)
di: Zhao, Ziyu, et al.
Pubblicazione: (2024)
TreeLoRA: Efficient Continual Learning via Layer-Wise LoRAs Guided by a Hierarchical Gradient-Similarity Tree
di: Qian, Yu-Yang, et al.
Pubblicazione: (2025)
di: Qian, Yu-Yang, et al.
Pubblicazione: (2025)
FlexLoRA: Entropy-Guided Flexible Low-Rank Adaptation
di: Liu, Muqing, et al.
Pubblicazione: (2026)
di: Liu, Muqing, et al.
Pubblicazione: (2026)
Restricted Block Permutation for Two-Sample Testing
di: Ho, Jungwoo
Pubblicazione: (2025)
di: Ho, Jungwoo
Pubblicazione: (2025)
Learning Permutation Distributions via Reflected Diffusion on Ranks
di: He, Sizhuang, et al.
Pubblicazione: (2026)
di: He, Sizhuang, et al.
Pubblicazione: (2026)
FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
di: Zou, Heming, et al.
Pubblicazione: (2025)
di: Zou, Heming, et al.
Pubblicazione: (2025)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
di: Li, Yuhang, et al.
Pubblicazione: (2024)
di: Li, Yuhang, et al.
Pubblicazione: (2024)
RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism
di: Sun, Mengyang, et al.
Pubblicazione: (2026)
di: Sun, Mengyang, et al.
Pubblicazione: (2026)
MetaLoRA: Tensor-Enhanced Adaptive Low-Rank Fine-tuning
di: Wang, Maolin, et al.
Pubblicazione: (2025)
di: Wang, Maolin, et al.
Pubblicazione: (2025)
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
di: Li, Mingjie, et al.
Pubblicazione: (2025)
di: Li, Mingjie, et al.
Pubblicazione: (2025)
RepLoRA: Reparameterizing Low-Rank Adaptation via the Perspective of Mixture of Experts
di: Truong, Tuan, et al.
Pubblicazione: (2025)
di: Truong, Tuan, et al.
Pubblicazione: (2025)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
di: Cai, Wen-Pu, et al.
Pubblicazione: (2024)
di: Cai, Wen-Pu, et al.
Pubblicazione: (2024)
LoRIF: Low-Rank Influence Functions for Scalable Training Data Attribution
di: Li, Shuangqi, et al.
Pubblicazione: (2026)
di: Li, Shuangqi, et al.
Pubblicazione: (2026)
Consensus Knowledge Graph Learning via Multi-view Sparse Low Rank Block Model
di: Cai, Tianxi, et al.
Pubblicazione: (2022)
di: Cai, Tianxi, et al.
Pubblicazione: (2022)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
di: Tao, Qian, et al.
Pubblicazione: (2024)
di: Tao, Qian, et al.
Pubblicazione: (2024)
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
di: Arai, Yamato, et al.
Pubblicazione: (2025)
di: Arai, Yamato, et al.
Pubblicazione: (2025)
LoRA+: Efficient Low Rank Adaptation of Large Models
di: Hayou, Soufiane, et al.
Pubblicazione: (2024)
di: Hayou, Soufiane, et al.
Pubblicazione: (2024)
Documenti analoghi
-
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026) -
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026) -
A method of using RSVD in residual calculation of LowBit GEMM
di: Gu, Hongyaoxing
Pubblicazione: (2024) -
PRoLoRA: Partial Rotation Empowers More Parameter-Efficient LoRA
di: Wang, Sheng, et al.
Pubblicazione: (2024) -
LoQT: Low-Rank Adapters for Quantized Pretraining
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)