RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ruicheng, Chen, Guangyu, Xu, Zunnan, Liu, Zihao, Zhong, Zhizhou, Zhang, Mingyang, Zhou, Jun, Li, Xiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026)
Identity-Consistent Video Generation under Large Facial-Angle Variations
von: Hu, Bin, et al.
Veröffentlicht: (2026)
von: Hu, Bin, et al.
Veröffentlicht: (2026)
RoboScape: Physics-informed Embodied World Model
von: Shang, Yu, et al.
Veröffentlicht: (2025)
von: Shang, Yu, et al.
Veröffentlicht: (2025)
Controllable Layer Decomposition for Reversible Multi-Layer Image Generation
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
Consistent123: One Image to Highly Consistent 3D Asset Using Case-Aware Diffusion Priors
von: Lin, Yukang, et al.
Veröffentlicht: (2023)
von: Lin, Yukang, et al.
Veröffentlicht: (2023)
TesserAct: Learning 4D Embodied World Models
von: Zhen, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2025)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction
von: Chan, Nga Teng, et al.
Veröffentlicht: (2026)
von: Chan, Nga Teng, et al.
Veröffentlicht: (2026)
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning
von: Xu, Peiran, et al.
Veröffentlicht: (2026)
von: Xu, Peiran, et al.
Veröffentlicht: (2026)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Embody4D: A Generalist 4D World Model for Embodied AI
von: Tu, Peiyan, et al.
Veröffentlicht: (2026)
von: Tu, Peiyan, et al.
Veröffentlicht: (2026)
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2025)
von: Hong, Fa-Ting, et al.
Veröffentlicht: (2025)
Unified Medical Image Segmentation with State Space Modeling Snake
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025)
StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception
von: Han, Evans, et al.
Veröffentlicht: (2026)
von: Han, Evans, et al.
Veröffentlicht: (2026)
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
von: Han, Haonan, et al.
Veröffentlicht: (2024)
von: Han, Haonan, et al.
Veröffentlicht: (2024)
RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
von: Sun, Xiaoquan, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoquan, et al.
Veröffentlicht: (2025)
RoboLayout: Differentiable 3D Scene Generation for Embodied Agents
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
von: Wu, Leyi, et al.
Veröffentlicht: (2026)
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
von: Shang, Yu, et al.
Veröffentlicht: (2026)
von: Shang, Yu, et al.
Veröffentlicht: (2026)
Divide-and-Conquer: Dual-Hierarchical Optimization for Semantic 4D Gaussian Spatting
von: Yan, Zhiying, et al.
Veröffentlicht: (2025)
von: Yan, Zhiying, et al.
Veröffentlicht: (2025)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
von: Guo, Xianda, et al.
Veröffentlicht: (2024)
RoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RL
von: Tang, Yinzhou, et al.
Veröffentlicht: (2025)
von: Tang, Yinzhou, et al.
Veröffentlicht: (2025)
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
von: Cheng, Junda, et al.
Veröffentlicht: (2025)
von: Cheng, Junda, et al.
Veröffentlicht: (2025)
StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors
von: Shen, Guibao, et al.
Veröffentlicht: (2025)
von: Shen, Guibao, et al.
Veröffentlicht: (2025)
MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models
von: Xu, Zunnan, et al.
Veröffentlicht: (2024)
von: Xu, Zunnan, et al.
Veröffentlicht: (2024)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
von: Xie, Yifan, et al.
Veröffentlicht: (2025)
von: Xie, Yifan, et al.
Veröffentlicht: (2025)
REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment
von: Han, Haonan, et al.
Veröffentlicht: (2024)
von: Han, Haonan, et al.
Veröffentlicht: (2024)
SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
von: Huang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Huang, Jiaqi, et al.
Veröffentlicht: (2025)
RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2025)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
4D Driving Scene Generation With Stereo Forcing
von: Lu, Hao, et al.
Veröffentlicht: (2025)
von: Lu, Hao, et al.
Veröffentlicht: (2025)
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
UniTT-Stereo: Unified Training of Transformer for Enhanced Stereo Matching
von: Kim, Soomin, et al.
Veröffentlicht: (2024)
von: Kim, Soomin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025) -
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2025) -
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
von: Zhang, Ruicheng, et al.
Veröffentlicht: (2026) -
Identity-Consistent Video Generation under Large Facial-Angle Variations
von: Hu, Bin, et al.
Veröffentlicht: (2026) -
RoboScape: Physics-informed Embodied World Model
von: Shang, Yu, et al.
Veröffentlicht: (2025)