Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Yuan, Bian, Yuexin, He, Sihong, Feng, Jie, Su, Qing, Han, Songyang, Petit, Jonathan, Ji, Shihao, Shi, Yuanyuan, Miao, Fei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917467573452800
author Zhuang, Yuan
Bian, Yuexin
He, Sihong
Feng, Jie
Su, Qing
Han, Songyang
Petit, Jonathan
Ji, Shihao
Shi, Yuanyuan
Miao, Fei
author_facet Zhuang, Yuan
Bian, Yuexin
He, Sihong
Feng, Jie
Su, Qing
Han, Songyang
Petit, Jonathan
Ji, Shihao
Shi, Yuanyuan
Miao, Fei
contents Scaling critic capacity is a promising direction for improving off-policy reinforcement learning (RL). However, recent work shows that larger critics are prone to overfitting and instability in replay-based bootstrapped training. In this paper, we propose using Low-Rank Adaptation (LoRA) as a structural regularizer for critic learning. Our approach freezes randomly initialized base matrices and optimizes only the corresponding low-rank adapters, thereby constraining critic updates to a low-dimensional subspace. We evaluate our method across different off-policy RL algorithms, including SAC and FastTD3 based on different network architectures. Empirically, LoRA efficiently reduces critic loss during training and improves overall policy performance, achieving the best or competitive results on most tasks. Extensive experiments demonstrate that our low-rank updates provide a simple and effective form of structural regularization for critic learning in off-policy RL.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18978
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning
Zhuang, Yuan
Bian, Yuexin
He, Sihong
Feng, Jie
Su, Qing
Han, Songyang
Petit, Jonathan
Ji, Shihao
Shi, Yuanyuan
Miao, Fei
Machine Learning
Artificial Intelligence
Scaling critic capacity is a promising direction for improving off-policy reinforcement learning (RL). However, recent work shows that larger critics are prone to overfitting and instability in replay-based bootstrapped training. In this paper, we propose using Low-Rank Adaptation (LoRA) as a structural regularizer for critic learning. Our approach freezes randomly initialized base matrices and optimizes only the corresponding low-rank adapters, thereby constraining critic updates to a low-dimensional subspace. We evaluate our method across different off-policy RL algorithms, including SAC and FastTD3 based on different network architectures. Empirically, LoRA efficiently reduces critic loss during training and improves overall policy performance, achieving the best or competitive results on most tasks. Extensive experiments demonstrate that our low-rank updates provide a simple and effective form of structural regularization for critic learning in off-policy RL.
title Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.18978