Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Hao, Li, Yuqi, Gao, Yuan, Xu, Fan, Zhang, Fan, Wang, Kun, Zhao, Penghao, Wang, Qiufeng, Zhao, Yizhou, Wang, Weiyan, Tian, Yingli, Wu, Xian, Huang, Xiaomeng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2605.03821
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913091123412992
author Wu, Hao
Li, Yuqi
Gao, Yuan
Xu, Fan
Zhang, Fan
Wang, Kun
Zhao, Penghao
Wang, Qiufeng
Zhao, Yizhou
Wang, Weiyan
Tian, Yingli
Wu, Xian
Huang, Xiaomeng
author_facet Wu, Hao
Li, Yuqi
Gao, Yuan
Xu, Fan
Zhang, Fan
Wang, Kun
Zhao, Penghao
Wang, Qiufeng
Zhao, Yizhou
Wang, Weiyan
Tian, Yingli
Wu, Xian
Huang, Xiaomeng
contents Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffer from error accumulation in long-horizon autoregressive prediction. We present RoboAlign-R1, a framework that combines reward-aligned post-training with stabilized long-horizon inference for robot video world models. We construct RobotWorldBench, a benchmark of 10,000 annotated video-instruction pairs collected from four robot data sources, and train a multimodal teacher judge, RoboAlign-Judge, to provide fine-grained six-dimensional evaluation of generated videos. We then distill the teacher into a lightweight student reward model for efficient reinforcement-learning-based post-training. To reduce long-horizon rollout drift, we further introduce Sliding Window Re-encoding (SWR), a training-free inference strategy that periodically refreshes the generation context. Under our in-domain evaluation protocol, RoboAlign-R1 improves the aggregate six-dimension score by 10.1% over the strongest baseline, including gains of 7.5% on Manipulation Accuracy and 4.6% on Instruction Following; these ranking improvements are further supported by an external VLM-based cross-check and a blinded human study. Meanwhile, SWR improves long-horizon prediction quality with only about 1% additional latency, yielding a 2.8% gain in SSIM and a 9.8% reduction in LPIPS. Together, these results show that reward-aligned post-training and stabilized long-horizon decoding improve task consistency, physical realism, and long-horizon prediction quality in robot video world models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
Wu, Hao
Li, Yuqi
Gao, Yuan
Xu, Fan
Zhang, Fan
Wang, Kun
Zhao, Penghao
Wang, Qiufeng
Zhao, Yizhou
Wang, Weiyan
Tian, Yingli
Wu, Xian
Huang, Xiaomeng
Robotics
Artificial Intelligence
Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffer from error accumulation in long-horizon autoregressive prediction. We present RoboAlign-R1, a framework that combines reward-aligned post-training with stabilized long-horizon inference for robot video world models. We construct RobotWorldBench, a benchmark of 10,000 annotated video-instruction pairs collected from four robot data sources, and train a multimodal teacher judge, RoboAlign-Judge, to provide fine-grained six-dimensional evaluation of generated videos. We then distill the teacher into a lightweight student reward model for efficient reinforcement-learning-based post-training. To reduce long-horizon rollout drift, we further introduce Sliding Window Re-encoding (SWR), a training-free inference strategy that periodically refreshes the generation context. Under our in-domain evaluation protocol, RoboAlign-R1 improves the aggregate six-dimension score by 10.1% over the strongest baseline, including gains of 7.5% on Manipulation Accuracy and 4.6% on Instruction Following; these ranking improvements are further supported by an external VLM-based cross-check and a blinded human study. Meanwhile, SWR improves long-horizon prediction quality with only about 1% additional latency, yielding a 2.8% gain in SSIM and a 9.8% reduction in LPIPS. Together, these results show that reward-aligned post-training and stabilized long-horizon decoding improve task consistency, physical realism, and long-horizon prediction quality in robot video world models.
title RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2605.03821