PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yang, Zhao, Jiangyuan, Fan, Chenyou, Yan, Fangzheng, Li, Tian, Tang, Haitong, Fu, Sen, Wu, Xuan'er, Weng, Qizhen, Zhang, Weinan, Li, Xiu, Zhang, Chi, Bai, Chenjia, Li, Xuelong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909003716493312
author Zhang, Yang
Zhao, Jiangyuan
Fan, Chenyou
Yan, Fangzheng
Li, Tian
Tang, Haitong
Fu, Sen
Wu, Xuan'er
Weng, Qizhen
Zhang, Weinan
Li, Xiu
Zhang, Chi
Bai, Chenjia
Li, Xuelong
author_facet Zhang, Yang
Zhao, Jiangyuan
Fan, Chenyou
Yan, Fangzheng
Li, Tian
Tang, Haitong
Fu, Sen
Wu, Xuan'er
Weng, Qizhen
Zhang, Weinan
Li, Xiu
Zhang, Chi
Bai, Chenjia
Li, Xuelong
contents Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloning, overlooking the fundamental nature of robot learning as a goal-reaching process that requires understanding temporal task progress. We present \textbf{PRTS} (\textbf{P}rimitive \textbf{R}easoning and \textbf{T}asking \textbf{S}ystem), a VLA foundation model that reformulates pretraining through Goal-Conditioned Reinforcement Learning. By treating language instructions as goals and employing contrastive reinforcement learning, PRTS learns a unified embedding space where the inner product of state-action and goal embeddings approximates the log-discounted goal occupancy, the probability of reaching the language-specified goal from the current state-action, quantitatively assessing physical feasibility beyond static semantic matching. PRTS draws this dense goal-reachability supervision directly from offline trajectories without reward annotations, and folds it into the VLM backbone via a role-aware causal mask, incurring negligible overhead over vanilla behavior cloning. This paradigm endows the high-level reasoning system with intrinsic goal reachability awareness, bridging semantic reasoning and temporal task progress, and further benefits goal-conditioned action prediction. Pretrained on 167B tokens of diverse manipulation and embodied-reasoning data, PRTS reaches state-of-the-art performance on LIBERO, LIBERO-Pro, LIBERO-Plus, SimplerEnv, and a real-world suite of 14 complex tasks, with particularly substantial gains on long-horizon, contact-rich, and zero-shot novel-instruction settings, confirming that injecting goal-reachability awareness significantly improves both execution success and long-horizon planning of general-purpose robotic foundation policies.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27472
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
Zhang, Yang
Zhao, Jiangyuan
Fan, Chenyou
Yan, Fangzheng
Li, Tian
Tang, Haitong
Fu, Sen
Wu, Xuan'er
Weng, Qizhen
Zhang, Weinan
Li, Xiu
Zhang, Chi
Bai, Chenjia
Li, Xuelong
Artificial Intelligence
Machine Learning
Robotics
Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloning, overlooking the fundamental nature of robot learning as a goal-reaching process that requires understanding temporal task progress. We present \textbf{PRTS} (\textbf{P}rimitive \textbf{R}easoning and \textbf{T}asking \textbf{S}ystem), a VLA foundation model that reformulates pretraining through Goal-Conditioned Reinforcement Learning. By treating language instructions as goals and employing contrastive reinforcement learning, PRTS learns a unified embedding space where the inner product of state-action and goal embeddings approximates the log-discounted goal occupancy, the probability of reaching the language-specified goal from the current state-action, quantitatively assessing physical feasibility beyond static semantic matching. PRTS draws this dense goal-reachability supervision directly from offline trajectories without reward annotations, and folds it into the VLM backbone via a role-aware causal mask, incurring negligible overhead over vanilla behavior cloning. This paradigm endows the high-level reasoning system with intrinsic goal reachability awareness, bridging semantic reasoning and temporal task progress, and further benefits goal-conditioned action prediction. Pretrained on 167B tokens of diverse manipulation and embodied-reasoning data, PRTS reaches state-of-the-art performance on LIBERO, LIBERO-Pro, LIBERO-Plus, SimplerEnv, and a real-world suite of 14 complex tasks, with particularly substantial gains on long-horizon, contact-rich, and zero-shot novel-instruction settings, confirming that injecting goal-reachability awareness significantly improves both execution success and long-horizon planning of general-purpose robotic foundation policies.
title PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
topic Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2604.27472