Stereo World Model: Camera-Guided Stereo Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yang-Tian, Huang, Zehuan, Niu, Yifan, Ma, Lin, Cao, Yan-Pei, Ma, Yuewen, Qi, Xiaojuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918395300020224
author Sun, Yang-Tian
Huang, Zehuan
Niu, Yifan
Ma, Lin
Cao, Yan-Pei
Ma, Yuewen
Qi, Xiaojuan
author_facet Sun, Yang-Tian
Huang, Zehuan
Niu, Yifan
Ma, Lin
Cao, Yan-Pei
Ma, Yuewen
Qi, Xiaojuan
contents We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry directly from disparity. To efficiently achieve consistent stereo generation, our approach introduces two key designs: (1) a unified camera-frame RoPE that augments latent tokens with camera-aware rotary positional encoding, enabling relative, view- and time-consistent conditioning while preserving pretrained video priors via a stable attention initialization; and (2) a stereo-aware attention decomposition that factors full 4D attention into 3D intra-view attention plus horizontal row attention, leveraging the epipolar prior to capture disparity-aligned correspondences with substantially lower compute. Across benchmarks, StereoWorld improves stereo consistency, disparity accuracy, and camera-motion fidelity over strong monocular-then-convert pipelines, achieving more than 3x faster generation with an additional 5% gain in viewpoint consistency. Beyond benchmarks, StereoWorld enables end-to-end binocular VR rendering without depth estimation or inpainting, enhances embodied policy learning through metric-scale depth grounding, and is compatible with long-video distillation for extended interactive stereo synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17375
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stereo World Model: Camera-Guided Stereo Video Generation
Sun, Yang-Tian
Huang, Zehuan
Niu, Yifan
Ma, Lin
Cao, Yan-Pei
Ma, Yuewen
Qi, Xiaojuan
Computer Vision and Pattern Recognition
We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry directly from disparity. To efficiently achieve consistent stereo generation, our approach introduces two key designs: (1) a unified camera-frame RoPE that augments latent tokens with camera-aware rotary positional encoding, enabling relative, view- and time-consistent conditioning while preserving pretrained video priors via a stable attention initialization; and (2) a stereo-aware attention decomposition that factors full 4D attention into 3D intra-view attention plus horizontal row attention, leveraging the epipolar prior to capture disparity-aligned correspondences with substantially lower compute. Across benchmarks, StereoWorld improves stereo consistency, disparity accuracy, and camera-motion fidelity over strong monocular-then-convert pipelines, achieving more than 3x faster generation with an additional 5% gain in viewpoint consistency. Beyond benchmarks, StereoWorld enables end-to-end binocular VR rendering without depth estimation or inpainting, enhances embodied policy learning through metric-scale depth grounding, and is compatible with long-video distillation for extended interactive stereo synthesis.
title Stereo World Model: Camera-Guided Stereo Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.17375