CRL-VLA: Continual Vision-Language-Action Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Qixin, Zhang, Shuo, Zhang, Hongyin, Wang, Renjie, Zhao, Han, Zhao, Libang, Li, Runze, Wang, Donglin, Huang, Chao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917244607397888
author Zeng, Qixin
Zhang, Shuo
Zhang, Hongyin
Wang, Renjie
Zhao, Han
Zhao, Libang
Li, Runze
Wang, Donglin
Huang, Chao
author_facet Zeng, Qixin
Zhang, Shuo
Zhang, Hongyin
Wang, Renjie
Zhao, Han
Zhao, Libang
Li, Runze
Wang, Donglin
Huang, Chao
contents Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation through environmental interaction. Thus, Continual Reinforcement Learning (CRL) is a promising pathway for deploying VLA models in lifelong robotic scenarios, yet balancing stability (retaining old skills) and plasticity (learning new ones) remains a formidable challenge for existing methods. We introduce CRL-VLA, a framework for continual post-training of VLA models with rigorous theoretical bounds. We derive a unified performance bound linking the stability-plasticity trade-off to goal-conditioned advantage magnitude, scaled by policy divergence. CRL-VLA resolves this dilemma via asymmetric regulation: constraining advantage magnitudes on prior tasks while enabling controlled growth on new tasks. This is realized through a simple but effective dual-critic architecture with novel Goal-Conditioned Value Formulation (GCVF), where a frozen critic anchors semantic consistency and a trainable estimator drives adaptation. Experiments on the LIBERO benchmark demonstrate that CRL-VLA effectively harmonizes these conflicting objectives, outperforming baselines in both anti-forgetting and forward adaptation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03445
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CRL-VLA: Continual Vision-Language-Action Learning
Zeng, Qixin
Zhang, Shuo
Zhang, Hongyin
Wang, Renjie
Zhao, Han
Zhao, Libang
Li, Runze
Wang, Donglin
Huang, Chao
Artificial Intelligence
Machine Learning
Robotics
Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Language-Action (VLA) models to master dexterous manipulation through environmental interaction. Thus, Continual Reinforcement Learning (CRL) is a promising pathway for deploying VLA models in lifelong robotic scenarios, yet balancing stability (retaining old skills) and plasticity (learning new ones) remains a formidable challenge for existing methods. We introduce CRL-VLA, a framework for continual post-training of VLA models with rigorous theoretical bounds. We derive a unified performance bound linking the stability-plasticity trade-off to goal-conditioned advantage magnitude, scaled by policy divergence. CRL-VLA resolves this dilemma via asymmetric regulation: constraining advantage magnitudes on prior tasks while enabling controlled growth on new tasks. This is realized through a simple but effective dual-critic architecture with novel Goal-Conditioned Value Formulation (GCVF), where a frozen critic anchors semantic consistency and a trainable estimator drives adaptation. Experiments on the LIBERO benchmark demonstrate that CRL-VLA effectively harmonizes these conflicting objectives, outperforming baselines in both anti-forgetting and forward adaptation.
title CRL-VLA: Continual Vision-Language-Action Learning
topic Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2602.03445