WonderJourney: Going from Anywhere to Everywhere

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Hong-Xing, Duan, Haoyi, Hur, Junhwa, Sargent, Kyle, Rubinstein, Michael, Freeman, William T., Cole, Forrester, Sun, Deqing, Snavely, Noah, Wu, Jiajun, Herrmann, Charles
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910407650705408
author Yu, Hong-Xing
Duan, Haoyi
Hur, Junhwa
Sargent, Kyle
Rubinstein, Michael
Freeman, William T.
Cole, Forrester
Sun, Deqing
Snavely, Noah
Wu, Jiajun
Herrmann, Charles
author_facet Yu, Hong-Xing
Duan, Haoyi
Hur, Junhwa
Sargent, Kyle
Rubinstein, Michael
Freeman, William T.
Cole, Forrester
Sun, Deqing
Snavely, Noah
Wu, Jiajun
Herrmann, Charles
contents We introduce WonderJourney, a modularized framework for perpetual 3D scene generation. Unlike prior work on view generation that focuses on a single type of scenes, we start at any user-provided location (by a text description or an image) and generate a journey through a long sequence of diverse yet coherently connected 3D scenes. We leverage an LLM to generate textual descriptions of the scenes in this journey, a text-driven point cloud generation pipeline to make a compelling and coherent sequence of 3D scenes, and a large VLM to verify the generated scenes. We show compelling, diverse visual results across various scene types and styles, forming imaginary "wonderjourneys". Project website: https://kovenyu.com/WonderJourney/
format Preprint
id arxiv_https___arxiv_org_abs_2312_03884
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle WonderJourney: Going from Anywhere to Everywhere
Yu, Hong-Xing
Duan, Haoyi
Hur, Junhwa
Sargent, Kyle
Rubinstein, Michael
Freeman, William T.
Cole, Forrester
Sun, Deqing
Snavely, Noah
Wu, Jiajun
Herrmann, Charles
Computer Vision and Pattern Recognition
Graphics
We introduce WonderJourney, a modularized framework for perpetual 3D scene generation. Unlike prior work on view generation that focuses on a single type of scenes, we start at any user-provided location (by a text description or an image) and generate a journey through a long sequence of diverse yet coherently connected 3D scenes. We leverage an LLM to generate textual descriptions of the scenes in this journey, a text-driven point cloud generation pipeline to make a compelling and coherent sequence of 3D scenes, and a large VLM to verify the generated scenes. We show compelling, diverse visual results across various scene types and styles, forming imaginary "wonderjourneys". Project website: https://kovenyu.com/WonderJourney/
title WonderJourney: Going from Anywhere to Everywhere
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2312.03884