RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seungwook, Shi, Yichun, Li, Kejie, Cho, Minsu, Wang, Peng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914063991177216
author Kim, Seungwook
Shi, Yichun
Li, Kejie
Cho, Minsu
Wang, Peng
author_facet Kim, Seungwook
Shi, Yichun
Li, Kejie
Cho, Minsu
Wang, Peng
contents Generating synthetic multi-view images from a text prompt is an essential bridge to generating synthetic 3D assets. In this work, we introduce RapidMV, a novel text-to-multi-view generative model that can produce 32 multi-view synthetic images in just around 5 seconds. In essence, we propose a novel spatio-angular latent space, encoding both the spatial appearance and angular viewpoint deviations into a single latent for improved efficiency and multi-view consistency. We achieve effective training of RapidMV by strategically decomposing our training process into multiple steps. We demonstrate that RapidMV outperforms existing methods in terms of consistency and latency, with competitive quality and text-image alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
Kim, Seungwook
Shi, Yichun
Li, Kejie
Cho, Minsu
Wang, Peng
Computer Vision and Pattern Recognition
Generating synthetic multi-view images from a text prompt is an essential bridge to generating synthetic 3D assets. In this work, we introduce RapidMV, a novel text-to-multi-view generative model that can produce 32 multi-view synthetic images in just around 5 seconds. In essence, we propose a novel spatio-angular latent space, encoding both the spatial appearance and angular viewpoint deviations into a single latent for improved efficiency and multi-view consistency. We achieve effective training of RapidMV by strategically decomposing our training process into multiple steps. We demonstrate that RapidMV outperforms existing methods in terms of consistency and latency, with competitive quality and text-image alignment.
title RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.24410