RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Geonho, Lee, Janghwan, Hong, Sukjin, Kim, Minsoo, Ahn, Euijai, Chang, Du-Seong, Choi, Jungwook |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
by: Lee, Janghwan, et al.
Published: (2023)
by: Lee, Janghwan, et al.
Published: (2023)
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
by: Lee, Janghwan, et al.
Published: (2024)
by: Lee, Janghwan, et al.
Published: (2024)
Understanding LoRA as Knowledge Memory: An Empirical Analysis
by: Back, Seungju, et al.
Published: (2026)
by: Back, Seungju, et al.
Published: (2026)
TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
by: Jeon, Jeimin, et al.
Published: (2026)
by: Jeon, Jeimin, et al.
Published: (2026)
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
by: Kwak, Hyunseok, et al.
Published: (2025)
by: Kwak, Hyunseok, et al.
Published: (2025)
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
by: Ahn, Beomjin, et al.
Published: (2026)
by: Ahn, Beomjin, et al.
Published: (2026)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
by: Park, Minho, et al.
Published: (2025)
by: Park, Minho, et al.
Published: (2025)
LoRA on the Go: Instance-level Dynamic LoRA Selection and Merging
by: Lee, Seungeon, et al.
Published: (2025)
by: Lee, Seungeon, et al.
Published: (2025)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
Rank-Accuracy Trade-off for LoRA: A Gradient-Flow Analysis
by: Rushka, Michael, et al.
Published: (2026)
by: Rushka, Michael, et al.
Published: (2026)
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
by: Lee, Jung Hyun, et al.
Published: (2026)
by: Lee, Jung Hyun, et al.
Published: (2026)
DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
by: Deng, Guanzhi, et al.
Published: (2026)
by: Deng, Guanzhi, et al.
Published: (2026)
FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA
by: Lee, Seanie, et al.
Published: (2025)
by: Lee, Seanie, et al.
Published: (2025)
Riemannian Optimization for LoRA on the Stiefel Manifold
by: Park, Juneyoung, et al.
Published: (2025)
by: Park, Juneyoung, et al.
Published: (2025)
A Note on LoRA
by: Fomenko, Vlad, et al.
Published: (2024)
by: Fomenko, Vlad, et al.
Published: (2024)
QuAILoRA: Quantization-Aware Initialization for LoRA
by: Lawton, Neal, et al.
Published: (2024)
by: Lawton, Neal, et al.
Published: (2024)
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization
by: Meng, Xiangdi, et al.
Published: (2024)
by: Meng, Xiangdi, et al.
Published: (2024)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
by: Kim, Minsoo, et al.
Published: (2024)
by: Kim, Minsoo, et al.
Published: (2024)
InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
by: Kim, Minsoo, et al.
Published: (2025)
by: Kim, Minsoo, et al.
Published: (2025)
Bayesian-LoRA: LoRA based Parameter Efficient Fine-Tuning using Optimal Quantization levels and Rank Values trough Differentiable Bayesian Gates
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA
by: Lee, Sangyoon, et al.
Published: (2026)
by: Lee, Sangyoon, et al.
Published: (2026)
R-LoRA: Randomized Multi-Head LoRA for Efficient Multi-Task Learning
by: Liu, Jinda, et al.
Published: (2025)
by: Liu, Jinda, et al.
Published: (2025)
ABM-LoRA: Activation Boundary Matching for Fast Convergence in Low-Rank Adaptation
by: Lee, Dongha, et al.
Published: (2025)
by: Lee, Dongha, et al.
Published: (2025)
Hybrid-LoRA: Bridging Full Fine-Tuning and Low-Rank Adaptation for Post-Training
by: Zhang, Chengqian, et al.
Published: (2026)
by: Zhang, Chengqian, et al.
Published: (2026)
GeLoRA: Geometric Adaptive Ranks For Efficient LoRA Fine-tuning
by: Ed-dib, Abdessalam, et al.
Published: (2024)
by: Ed-dib, Abdessalam, et al.
Published: (2024)
A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search
by: Seong-Eun, Baek, et al.
Published: (2026)
by: Seong-Eun, Baek, et al.
Published: (2026)
Rethinking the Rank Threshold for LoRA Fine-Tuning
by: Park, Juneyoung
Published: (2026)
by: Park, Juneyoung
Published: (2026)
Post-Optimization Adaptive Rank Allocation for LoRA
by: Kumaravelu, Vishnuprasadh, et al.
Published: (2026)
by: Kumaravelu, Vishnuprasadh, et al.
Published: (2026)
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
by: Choi, Kanghyun, et al.
Published: (2025)
by: Choi, Kanghyun, et al.
Published: (2025)
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment
by: Fan, Chenghao, et al.
Published: (2025)
by: Fan, Chenghao, et al.
Published: (2025)
FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimation
by: Sathyavageeswaran, Ramakrishnan
Published: (2026)
by: Sathyavageeswaran, Ramakrishnan
Published: (2026)
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation
by: Hwang, Injoon, et al.
Published: (2024)
by: Hwang, Injoon, et al.
Published: (2024)
LoRA-Key: User-Centric LoRA Watermarking for Text-to-Image Diffusion Models
by: Wang, Yaopeng, et al.
Published: (2026)
by: Wang, Yaopeng, et al.
Published: (2026)
Run LoRA Run: Faster and Lighter LoRA Implementations
by: Cherniuk, Daria, et al.
Published: (2023)
by: Cherniuk, Daria, et al.
Published: (2023)
Dual LoRA: Enhancing LoRA with Magnitude and Direction Updates
by: Xu, Yixing, et al.
Published: (2025)
by: Xu, Yixing, et al.
Published: (2025)
Selectively Dilated Convolution for Accuracy-Preserving Sparse Pillar-based Embedded 3D Object Detection
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
Co-LoRA: Collaborative Model Personalization on Heterogeneous Multi-Modal Clients
by: Seo, Minhyuk, et al.
Published: (2025)
by: Seo, Minhyuk, et al.
Published: (2025)
LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
by: Che, Chang, et al.
Published: (2025)
by: Che, Chang, et al.
Published: (2025)
Similar Items
-
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
by: Lee, Janghwan, et al.
Published: (2024) -
Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization
by: Lee, Janghwan, et al.
Published: (2023) -
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
by: Lee, Janghwan, et al.
Published: (2024) -
Understanding LoRA as Knowledge Memory: An Empirical Analysis
by: Back, Seungju, et al.
Published: (2026) -
TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts
by: Jeon, Jeimin, et al.
Published: (2026)