The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Hanlin, Ouyang, Hao, Wang, Qiuyu, Yu, Yue, Meng, Yihao, Wang, Wen, Cheng, Ka Leong, Ma, Shuailei, Bai, Qingyan, Li, Yixuan, Chen, Cheng, Zeng, Yanhong, Zhu, Xing, Shen, Yujun, Chen, Qifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914208322420736
author Wang, Hanlin
Ouyang, Hao
Wang, Qiuyu
Yu, Yue
Meng, Yihao
Wang, Wen
Cheng, Ka Leong
Ma, Shuailei
Bai, Qingyan
Li, Yixuan
Chen, Cheng
Zeng, Yanhong
Zhu, Xing
Shen, Yujun
Chen, Qifeng
author_facet Wang, Hanlin
Ouyang, Hao
Wang, Qiuyu
Yu, Yue
Meng, Yihao
Wang, Wen
Cheng, Ka Leong
Ma, Shuailei
Bai, Qingyan
Li, Yixuan
Chen, Cheng
Zeng, Yanhong
Zhu, Xing
Shen, Yujun
Chen, Qifeng
contents We present WorldCanvas, a framework for promptable world events that enables rich, user-directed simulation by combining text, trajectories, and reference images. Unlike text-only approaches and existing trajectory-controlled image-to-video methods, our multimodal approach combines trajectories -- encoding motion, timing, and visibility -- with natural language for semantic intent and reference images for visual grounding of object identity, enabling the generation of coherent, controllable events that include multi-agent interactions, object entry/exit, reference-guided appearance and counterintuitive events. The resulting videos demonstrate not only temporal coherence but also emergent consistency, preserving object identity and scene despite temporary disappearance. By supporting expressive world events generation, WorldCanvas advances world models from passive predictors to interactive, user-shaped simulators. Our project page is available at: https://worldcanvas.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
Wang, Hanlin
Ouyang, Hao
Wang, Qiuyu
Yu, Yue
Meng, Yihao
Wang, Wen
Cheng, Ka Leong
Ma, Shuailei
Bai, Qingyan
Li, Yixuan
Chen, Cheng
Zeng, Yanhong
Zhu, Xing
Shen, Yujun
Chen, Qifeng
Computer Vision and Pattern Recognition
We present WorldCanvas, a framework for promptable world events that enables rich, user-directed simulation by combining text, trajectories, and reference images. Unlike text-only approaches and existing trajectory-controlled image-to-video methods, our multimodal approach combines trajectories -- encoding motion, timing, and visibility -- with natural language for semantic intent and reference images for visual grounding of object identity, enabling the generation of coherent, controllable events that include multi-agent interactions, object entry/exit, reference-guided appearance and counterintuitive events. The resulting videos demonstrate not only temporal coherence but also emergent consistency, preserving object identity and scene despite temporary disappearance. By supporting expressive world events generation, WorldCanvas advances world models from passive predictors to interactive, user-shaped simulators. Our project page is available at: https://worldcanvas.github.io/.
title The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.16924