A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhai, Shaopeng, Zhang, Qi, Zhang, Tianyi, Huang, Fuxian, Zhang, Haoran, Zhou, Ming, Zhang, Shengzhe, Liu, Litao, Lin, Sixu, Pang, Jiangmiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
by: Lin, Sixu, et al.
Published: (2026)
by: Lin, Sixu, et al.
Published: (2026)
CLSP: High-Fidelity Contrastive Language-State Pre-training for Agent State Representation
by: Huang, Fuxian, et al.
Published: (2024)
by: Huang, Fuxian, et al.
Published: (2024)
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
by: Huang, Alex S., et al.
Published: (2026)
by: Huang, Alex S., et al.
Published: (2026)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)
by: Zhang, Yihao, et al.
Published: (2025)
UniCon: A Unified System for Efficient Robot Learning Transfers
by: Lin, Yunfeng, et al.
Published: (2026)
by: Lin, Yunfeng, et al.
Published: (2026)
Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models
by: Wen, Siqi, et al.
Published: (2026)
by: Wen, Siqi, et al.
Published: (2026)
Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-Shot Real-World Deployment
by: Yan, Ningyu, et al.
Published: (2026)
by: Yan, Ningyu, et al.
Published: (2026)
Dexterous Grasping with Real-World Robotic Reinforcement Learning
by: Huang, Dongchi, et al.
Published: (2025)
by: Huang, Dongchi, et al.
Published: (2025)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
by: Xu, Xiaoxu, et al.
Published: (2026)
by: Xu, Xiaoxu, et al.
Published: (2026)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
RoboStriker: Hierarchical Decision-Making for Autonomous Humanoid Boxing
by: Yin, Kangning, et al.
Published: (2026)
by: Yin, Kangning, et al.
Published: (2026)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
MASQ: Multi-Agent Reinforcement Learning for Single Quadruped Robot Locomotion
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
exUMI: Extensible Robot Teaching System with Action-aware Task-agnostic Tactile Representation
by: Xu, Yue, et al.
Published: (2025)
by: Xu, Yue, et al.
Published: (2025)
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
by: Zhang, Zhilong, et al.
Published: (2026)
by: Zhang, Zhilong, et al.
Published: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
Building Open-Ended Embodied Agent via Language-Policy Bidirectional Adaptation
by: Zhai, Shaopeng, et al.
Published: (2023)
by: Zhai, Shaopeng, et al.
Published: (2023)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
by: Lv, Qi, et al.
Published: (2025)
by: Lv, Qi, et al.
Published: (2025)
MineRobot: A Unified Framework for Kinematics Modeling and Solving of Underground Mining Robots in Virtual Environments
by: Hou, Shengzhe, et al.
Published: (2026)
by: Hou, Shengzhe, et al.
Published: (2026)
H-Zero: Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer
by: Lin, Yunfeng, et al.
Published: (2025)
by: Lin, Yunfeng, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
RoboKeyGen: Robot Pose and Joint Angles Estimation via Diffusion-based 3D Keypoint Generation
by: Tian, Yang, et al.
Published: (2024)
by: Tian, Yang, et al.
Published: (2024)
Survey of Vision-Language-Action Models for Embodied Manipulation
by: Li, Haoran, et al.
Published: (2025)
by: Li, Haoran, et al.
Published: (2025)
Cooperative-Competitive Team Play of Real-World Craft Robots
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
Reinforcing Action Policies by Prophesying
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
by: Han, Xiaoshen, et al.
Published: (2025)
by: Han, Xiaoshen, et al.
Published: (2025)
TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation
by: Xu, Qinwen, et al.
Published: (2026)
by: Xu, Qinwen, et al.
Published: (2026)
Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
by: Cai, Rui, et al.
Published: (2026)
by: Cai, Rui, et al.
Published: (2026)
CO-RFT: Efficient Fine-Tuning of Vision-Language-Action Models through Chunked Offline Reinforcement Learning
by: Huang, Dongchi, et al.
Published: (2025)
by: Huang, Dongchi, et al.
Published: (2025)
TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments
by: Huang, Zhiyu, et al.
Published: (2026)
by: Huang, Zhiyu, et al.
Published: (2026)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
by: Zhang, Yichi, et al.
Published: (2026)
by: Zhang, Yichi, et al.
Published: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
by: Li, Runze, et al.
Published: (2026)
by: Li, Runze, et al.
Published: (2026)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
Similar Items
-
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization
by: Lin, Sixu, et al.
Published: (2026) -
CLSP: High-Fidelity Contrastive Language-State Pre-training for Agent State Representation
by: Huang, Fuxian, et al.
Published: (2024) -
Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy
by: Zhang, Tianyi, et al.
Published: (2025) -
VLA-REPLICA: A Low-Cost, Reproducible Benchmark for Real-World Evaluation of Vision-Language-Action Models
by: Huang, Alex S., et al.
Published: (2026) -
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
by: Zhang, Yihao, et al.
Published: (2025)