CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhai, Xuanran, Ou, Binkai, Yu, Qiaojun, Hao, Ce, Liu, Yaohua
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918320384507904
author Zhai, Xuanran
Ou, Binkai
Yu, Qiaojun
Hao, Ce
Liu, Yaohua
author_facet Zhai, Xuanran
Ou, Binkai
Yu, Qiaojun
Hao, Ce
Liu, Yaohua
contents Vision Language Action (VLA) models enable instruction following manipulation, yet dualarm deployment remains unsafe due to under modeled selfcollisions between arms and grasped objects. We introduce CoFreeVLA, which augments an endtoend VLA with a short horizon selfcollision risk estimator that predicts collision likelihood from proprioception, visual embeddings, and planned actions. The estimator gates risky commands, recovers to safe states via risk-guided adjustments, and shapes policy refinement for safer rollouts. It is pre-trained with model-based collision labels and posttrained on real robot rollouts for calibration. On five bimanual tasks with the PiPER robot arm, CoFreeVLA reduces selfcollisions and improves success rates versus RDT and APEX.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21712
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation
Zhai, Xuanran
Ou, Binkai
Yu, Qiaojun
Hao, Ce
Liu, Yaohua
Robotics
Vision Language Action (VLA) models enable instruction following manipulation, yet dualarm deployment remains unsafe due to under modeled selfcollisions between arms and grasped objects. We introduce CoFreeVLA, which augments an endtoend VLA with a short horizon selfcollision risk estimator that predicts collision likelihood from proprioception, visual embeddings, and planned actions. The estimator gates risky commands, recovers to safe states via risk-guided adjustments, and shapes policy refinement for safer rollouts. It is pre-trained with model-based collision labels and posttrained on real robot rollouts for calibration. On five bimanual tasks with the PiPER robot arm, CoFreeVLA reduces selfcollisions and improves success rates versus RDT and APEX.
title CoFreeVLA: Collision-Free Dual-Arm Manipulation via Vision-Language-Action Model and Risk Estimation
topic Robotics
url https://arxiv.org/abs/2601.21712