ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Junjie, Dong, Jiahao, Wang, Yingheng, De Sa, Christopher, Kuleshov, Volodymyr |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
by: Tseng, Albert, et al.
Published: (2024)
by: Tseng, Albert, et al.
Published: (2024)
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
by: Chee, Jerry, et al.
Published: (2023)
by: Chee, Jerry, et al.
Published: (2023)
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
by: Avinash, MSR
Published: (2025)
by: Avinash, MSR
Published: (2025)
Robust Federated Finetuning of LLMs via Alternating Optimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2025)
by: Chen, Shuangyi, et al.
Published: (2025)
The Impact of Initialization on LoRA Finetuning Dynamics
by: Hayou, Soufiane, et al.
Published: (2024)
by: Hayou, Soufiane, et al.
Published: (2024)
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization
by: Deng, Yanxia, et al.
Published: (2025)
by: Deng, Yanxia, et al.
Published: (2025)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs
by: Ye, Zhengmao, et al.
Published: (2023)
by: Ye, Zhengmao, et al.
Published: (2023)
ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
by: Arabpour, Reza, et al.
Published: (2025)
by: Arabpour, Reza, et al.
Published: (2025)
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing
by: Li, Wenbing, et al.
Published: (2025)
by: Li, Wenbing, et al.
Published: (2025)
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
by: Zhao, Zhixiong, et al.
Published: (2025)
by: Zhao, Zhixiong, et al.
Published: (2025)
$\textit{Trans-LoRA}$: towards data-free Transferable Parameter Efficient Finetuning
by: Wang, Runqian, et al.
Published: (2024)
by: Wang, Runqian, et al.
Published: (2024)
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
by: Chen, Nan, et al.
Published: (2026)
by: Chen, Nan, et al.
Published: (2026)
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
by: Choi, Kanghyun, et al.
Published: (2025)
by: Choi, Kanghyun, et al.
Published: (2025)
NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs
by: Wang, Shuaidi, et al.
Published: (2026)
by: Wang, Shuaidi, et al.
Published: (2026)
Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
by: Liu, Jiahao, et al.
Published: (2025)
by: Liu, Jiahao, et al.
Published: (2025)
Bayesian-LoRA: LoRA based Parameter Efficient Fine-Tuning using Optimal Quantization levels and Rank Values trough Differentiable Bayesian Gates
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMs
by: Lee, Jung Hyun, et al.
Published: (2026)
by: Lee, Jung Hyun, et al.
Published: (2026)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
by: Huang, Xijie, et al.
Published: (2024)
by: Huang, Xijie, et al.
Published: (2024)
FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets
by: Wang, Dannong, et al.
Published: (2025)
by: Wang, Dannong, et al.
Published: (2025)
Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantization
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
Activated LoRA: Fine-tuned LLMs for Intrinsics
by: Greenewald, Kristjan, et al.
Published: (2025)
by: Greenewald, Kristjan, et al.
Published: (2025)
LoRA is All You Need for Safety Alignment of Reasoning LLMs
by: Xue, Yihao, et al.
Published: (2025)
by: Xue, Yihao, et al.
Published: (2025)
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning
by: Xu, Ziwen, et al.
Published: (2026)
by: Xu, Ziwen, et al.
Published: (2026)
HBLLM: Wavelet-Enhanced High-Fidelity 1-Bit Quantization for LLMs
by: Chen, Ningning, et al.
Published: (2025)
by: Chen, Ningning, et al.
Published: (2025)
Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
by: Ao, Shuang, et al.
Published: (2025)
by: Ao, Shuang, et al.
Published: (2025)
RanDeS: Randomized Delta Superposition for Multi-Model Compression
by: Zhou, Hangyu, et al.
Published: (2025)
by: Zhou, Hangyu, et al.
Published: (2025)
LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
by: Salles, Marcel Mateos, et al.
Published: (2025)
by: Salles, Marcel Mateos, et al.
Published: (2025)
Active Preference Inference using Language Models and Probabilistic Reasoning
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)
by: Piriyakulkij, Wasu Top, et al.
Published: (2023)
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
by: Zhu, Zhanda, et al.
Published: (2025)
by: Zhu, Zhanda, et al.
Published: (2025)
Fast and Continual Knowledge Graph Embedding via Incremental LoRA
by: Liu, Jiajun, et al.
Published: (2024)
by: Liu, Jiajun, et al.
Published: (2024)
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
by: Sheng, Ying, et al.
Published: (2023)
by: Sheng, Ying, et al.
Published: (2023)
Parameter-Efficient Fine-Tuning for HAR: Integrating LoRA and QLoRA into Transformer Models
by: Seregina, Irina, et al.
Published: (2025)
by: Seregina, Irina, et al.
Published: (2025)
ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMs
by: Chen, Xunlei, et al.
Published: (2026)
by: Chen, Xunlei, et al.
Published: (2026)
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-Tuning
by: Tian, Chunlin, et al.
Published: (2024)
by: Tian, Chunlin, et al.
Published: (2024)
FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance
by: Chung, Hyunsuk, et al.
Published: (2026)
by: Chung, Hyunsuk, et al.
Published: (2026)
SDXL Finetuned with LoRA for Coloring Therapy: Generating Graphic Templates Inspired by United Arab Emirates Culture
by: Alfalasi, Abdulla, et al.
Published: (2024)
by: Alfalasi, Abdulla, et al.
Published: (2024)
Similar Items
-
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
by: Tseng, Albert, et al.
Published: (2024) -
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
by: Chee, Jerry, et al.
Published: (2023) -
SplitQuantV2: Enhancing Low-Bit Quantization of LLMs Without GPUs
by: Song, Jaewoo, et al.
Published: (2025) -
Profiling LoRA/QLoRA Fine-Tuning Efficiency on Consumer GPUs: An RTX 4060 Case Study
by: Avinash, MSR
Published: (2025) -
Robust Federated Finetuning of LLMs via Alternating Optimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2025)