DiffLoRA: Differential Low-Rank Adapters for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Misrahi, Alexandre, Chirkova, Nadezhda, Louis, Maxime, Nikoulina, Vassilina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915419970863104
author Misrahi, Alexandre
Chirkova, Nadezhda
Louis, Maxime
Nikoulina, Vassilina
author_facet Misrahi, Alexandre
Chirkova, Nadezhda
Louis, Maxime
Nikoulina, Vassilina
contents Differential Transformer has recently been proposed to improve performance in Transformer models by canceling out noise through a denoiser attention mechanism. In this work, we introduce DiffLoRA, a parameter-efficient adaptation of the differential attention mechanism, with low-rank adapters on both positive and negative attention terms. This approach retains the efficiency of LoRA while aiming to benefit from the performance gains of differential attention. We evaluate DiffLoRA across a broad range of NLP tasks, including general benchmarks, many-shot in-context learning, RAG, and long-context tests. We observe that, although DiffLoRA falls short of other parameter-efficient fine-tuning methods in most evaluation tasks, it shows interesting results in certain domains (+11 pts on LoRA for HumanEval). We analyze the attention patterns post-finetuning to identify the reasons for this behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23588
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffLoRA: Differential Low-Rank Adapters for Large Language Models
Misrahi, Alexandre
Chirkova, Nadezhda
Louis, Maxime
Nikoulina, Vassilina
Computation and Language
Differential Transformer has recently been proposed to improve performance in Transformer models by canceling out noise through a denoiser attention mechanism. In this work, we introduce DiffLoRA, a parameter-efficient adaptation of the differential attention mechanism, with low-rank adapters on both positive and negative attention terms. This approach retains the efficiency of LoRA while aiming to benefit from the performance gains of differential attention. We evaluate DiffLoRA across a broad range of NLP tasks, including general benchmarks, many-shot in-context learning, RAG, and long-context tests. We observe that, although DiffLoRA falls short of other parameter-efficient fine-tuning methods in most evaluation tasks, it shows interesting results in certain domains (+11 pts on LoRA for HumanEval). We analyze the attention patterns post-finetuning to identify the reasons for this behavior.
title DiffLoRA: Differential Low-Rank Adapters for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.23588