Mask World Model: Predicting What Matters for Robust Robot Policy Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lou, Yunfan, Chi, Xiaowei, Zhang, Xiaojie, Qian, Zezhong, Li, Chengxuan, Zhang, Rongyu, Lyu, Yaoxu, Song, Guoyu, Fu, Chuyao, Xu, Haoxuan, Wang, Pengwei, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
von: Qian, Zezhong, et al.
Veröffentlicht: (2025)
von: Qian, Zezhong, et al.
Veröffentlicht: (2025)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
von: Li, Guangrun, et al.
Veröffentlicht: (2025)
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
EVA: An Embodied World Model for Future Video Anticipation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets
von: Wu, Zhuangzhe, et al.
Veröffentlicht: (2026)
von: Wu, Zhuangzhe, et al.
Veröffentlicht: (2026)
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
BEVUDA: Multi-geometric Space Alignments for Domain Adaptive BEV 3D Object Detection
von: Liu, Jiaming, et al.
Veröffentlicht: (2022)
von: Liu, Jiaming, et al.
Veröffentlicht: (2022)
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
Can World Models Benefit VLMs for World Dynamics?
von: Zhang, Kevin, et al.
Veröffentlicht: (2025)
von: Zhang, Kevin, et al.
Veröffentlicht: (2025)
RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance
von: Li, Ying, et al.
Veröffentlicht: (2025)
von: Li, Ying, et al.
Veröffentlicht: (2025)
SpikeGen: Decoupled "Rods and Cones" Visual Representation Processing with Latent Generative Framework
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
von: Han, Yi, et al.
Veröffentlicht: (2025)
von: Han, Yi, et al.
Veröffentlicht: (2025)
Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
von: Jia, Yueru, et al.
Veröffentlicht: (2025)
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
von: Tan, Huajie, et al.
Veröffentlicht: (2025)
ProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-Evolution
von: Fu, Chuyao, et al.
Veröffentlicht: (2026)
von: Fu, Chuyao, et al.
Veröffentlicht: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
von: Liu, Mengzhen, et al.
Veröffentlicht: (2026)
Real-Time Robot Execution with Masked Action Chunking
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
von: Wang, Haoxuan, et al.
Veröffentlicht: (2026)
OmniSAT: Compact Action Token, Faster Auto Regression
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
Mean Masked Autoencoder with Flow-Mixing for Encrypted Traffic Classification
von: Liu, Xiao, et al.
Veröffentlicht: (2026)
von: Liu, Xiao, et al.
Veröffentlicht: (2026)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
von: Ni, Ziqi, et al.
Veröffentlicht: (2025)
von: Ni, Ziqi, et al.
Veröffentlicht: (2025)
Latent Anomaly Detection: Masked VQ-GAN for Unsupervised Segmentation in Medical CBCT
von: Wang, Pengwei
Veröffentlicht: (2025)
von: Wang, Pengwei
Veröffentlicht: (2025)
Automated Genomic Interpretation via Concept Bottleneck Models for Medical Robotics
von: Li, Zijun, et al.
Veröffentlicht: (2025)
von: Li, Zijun, et al.
Veröffentlicht: (2025)
Attention Masks Help Adversarial Attacks to Bypass Safety Detectors
von: Shi, Yunfan
Veröffentlicht: (2024)
von: Shi, Yunfan
Veröffentlicht: (2024)
HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
von: Feng, Qiuxuan, et al.
Veröffentlicht: (2026)
von: Feng, Qiuxuan, et al.
Veröffentlicht: (2026)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
von: Tan, Huajie, et al.
Veröffentlicht: (2026)
von: Tan, Huajie, et al.
Veröffentlicht: (2026)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
von: Zhou, Enshen, et al.
Veröffentlicht: (2025)
Proxy Robustness in Vision Language Models is Effortlessly Transferable
von: Fu, Xiaowei, et al.
Veröffentlicht: (2026)
von: Fu, Xiaowei, et al.
Veröffentlicht: (2026)
Towards Robust Image Denoising with Scale Equivariance
von: Zhang, Dawei, et al.
Veröffentlicht: (2025)
von: Zhang, Dawei, et al.
Veröffentlicht: (2025)
TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion
von: Mi, Weishi, et al.
Veröffentlicht: (2026)
von: Mi, Weishi, et al.
Veröffentlicht: (2026)
T-REX: Mixture-of-Rank-One-Experts with Semantic-aware Intuition for Multi-task Large Language Model Finetuning
von: Zhang, Rongyu, et al.
Veröffentlicht: (2024)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2024)
RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
von: Tang, Yingbo, et al.
Veröffentlicht: (2025)
von: Tang, Yingbo, et al.
Veröffentlicht: (2025)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
von: Peng, Cheng, et al.
Veröffentlicht: (2025)
von: Peng, Cheng, et al.
Veröffentlicht: (2025)
Adaptive Label Correction for Robust Medical Image Segmentation with Noisy Labels
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
Orochi: Versatile Biomedical Image Processor
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
von: Qian, Zezhong, et al.
Veröffentlicht: (2025) -
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
von: Li, Guangrun, et al.
Veröffentlicht: (2025) -
M$^{2}$Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2023) -
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025) -
EVA: An Embodied World Model for Future Video Anticipation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)