Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Moule, Guan, Shuhao, Patane, Andrea, Gregg, David, Botterweck, Goetz
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908965633261568
author Lin, Moule
Guan, Shuhao
Patane, Andrea
Gregg, David
Botterweck, Goetz
author_facet Lin, Moule
Guan, Shuhao
Patane, Andrea
Gregg, David
Botterweck, Goetz
contents Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration. In this work, we introduce Bayesian-LoRA, which reformulates the deterministic LoRA update as a probabilistic low-rank representation inspired by Sparse Gaussian Processes. We identify a structural isomorphism between LoRA's factorization and Kronecker-factored SGP posteriors, and show that LoRA emerges as a limiting case when posterior uncertainty collapses. We conduct extensive experiments on various LLM architectures across commonsense reasoning benchmarks. With only approximately 0.42M additional parameters and ${\approx}1.2{\times}$ training cost relative to standard LoRA, Bayesian-LoRA significantly improves calibration across models up to 30B, achieving up to 84% ECE reduction and 76% NLL reduction while maintaining competitive accuracy for both in-distribution and out-of-distribution (OoD) evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21003
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
Lin, Moule
Guan, Shuhao
Patane, Andrea
Gregg, David
Botterweck, Goetz
Artificial Intelligence
Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration. In this work, we introduce Bayesian-LoRA, which reformulates the deterministic LoRA update as a probabilistic low-rank representation inspired by Sparse Gaussian Processes. We identify a structural isomorphism between LoRA's factorization and Kronecker-factored SGP posteriors, and show that LoRA emerges as a limiting case when posterior uncertainty collapses. We conduct extensive experiments on various LLM architectures across commonsense reasoning benchmarks. With only approximately 0.42M additional parameters and ${\approx}1.2{\times}$ training cost relative to standard LoRA, Bayesian-LoRA significantly improves calibration across models up to 30B, achieving up to 84% ECE reduction and 76% NLL reduction while maintaining competitive accuracy for both in-distribution and out-of-distribution (OoD) evaluations.
title Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2601.21003