RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yuxuan, Li, Xiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912443872051200
author Chen, Yuxuan
Li, Xiao
author_facet Chen, Yuxuan
Li, Xiao
contents Vision-Language-Action models (VLA) have demonstrated remarkable capabilities and promising potential in solving complex robotic manipulation tasks. However, their substantial parameter sizes and high inference latency pose significant challenges for real-world deployment, particularly on resource-constrained robotic platforms. To address this issue, we begin by conducting an extensive empirical study to explore the effectiveness of model compression techniques when applied to VLAs. Building on the insights gained from these preliminary experiments, we propose RLRC, a three-stage recovery method for compressed VLAs, including structured pruning, performance recovery based on SFT and RL, and further quantization. RLRC achieves up to an 8x reduction in memory usage and a 2.3x improvement in inference throughput, while maintaining or even surpassing the original VLA's task success rate. Extensive experiments show that RLRC consistently outperforms existing compression baselines, demonstrating strong potential for on-device deployment of VLAs. Project website: https://rlrc-vla.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2506_17639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
Chen, Yuxuan
Li, Xiao
Robotics
Artificial Intelligence
Vision-Language-Action models (VLA) have demonstrated remarkable capabilities and promising potential in solving complex robotic manipulation tasks. However, their substantial parameter sizes and high inference latency pose significant challenges for real-world deployment, particularly on resource-constrained robotic platforms. To address this issue, we begin by conducting an extensive empirical study to explore the effectiveness of model compression techniques when applied to VLAs. Building on the insights gained from these preliminary experiments, we propose RLRC, a three-stage recovery method for compressed VLAs, including structured pruning, performance recovery based on SFT and RL, and further quantization. RLRC achieves up to an 8x reduction in memory usage and a 2.3x improvement in inference throughput, while maintaining or even surpassing the original VLA's task success rate. Extensive experiments show that RLRC consistently outperforms existing compression baselines, demonstrating strong potential for on-device deployment of VLAs. Project website: https://rlrc-vla.github.io
title RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2506.17639