Improving LoRA with Variational Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cong, Bai, Daheim, Nico, Shen, Yuesong, Yokota, Rio, Khan, Mohammad Emtiyaz, Möllenhoff, Thomas
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909651429228544
author Cong, Bai
Daheim, Nico
Shen, Yuesong
Yokota, Rio
Khan, Mohammad Emtiyaz
Möllenhoff, Thomas
author_facet Cong, Bai
Daheim, Nico
Shen, Yuesong
Yokota, Rio
Khan, Mohammad Emtiyaz
Möllenhoff, Thomas
contents Bayesian methods have recently been used to improve LoRA finetuning and, although they improve calibration, their effect on other metrics (such as accuracy) is marginal and can sometimes even be detrimental. Moreover, Bayesian methods also increase computational overheads and require additional tricks for them to work well. Here, we fix these issues by using a recently proposed variational algorithm called IVON. We show that IVON is easy to implement and has similar costs to AdamW, and yet it can also drastically improve many metrics by using a simple posterior pruning technique. We present extensive results on billion-scale LLMs (Llama and Qwen series) going way beyond the scale of existing applications of IVON. For example, we finetune a Llama-3.2-3B model on a set of commonsense reasoning tasks and improve accuracy over AdamW by 1.3% and reduce ECE by 5.4%, outperforming AdamW and other recent Bayesian methods like Laplace-LoRA and BLoB. Overall, our results show that variational learning with IVON can effectively improve LoRA finetuning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving LoRA with Variational Learning
Cong, Bai
Daheim, Nico
Shen, Yuesong
Yokota, Rio
Khan, Mohammad Emtiyaz
Möllenhoff, Thomas
Machine Learning
Artificial Intelligence
Computation and Language
Bayesian methods have recently been used to improve LoRA finetuning and, although they improve calibration, their effect on other metrics (such as accuracy) is marginal and can sometimes even be detrimental. Moreover, Bayesian methods also increase computational overheads and require additional tricks for them to work well. Here, we fix these issues by using a recently proposed variational algorithm called IVON. We show that IVON is easy to implement and has similar costs to AdamW, and yet it can also drastically improve many metrics by using a simple posterior pruning technique. We present extensive results on billion-scale LLMs (Llama and Qwen series) going way beyond the scale of existing applications of IVON. For example, we finetune a Llama-3.2-3B model on a set of commonsense reasoning tasks and improve accuracy over AdamW by 1.3% and reduce ECE by 5.4%, outperforming AdamW and other recent Bayesian methods like Laplace-LoRA and BLoB. Overall, our results show that variational learning with IVON can effectively improve LoRA finetuning.
title Improving LoRA with Variational Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.14280