FlexGen: Flexible Multi-View Generation from Text and Image Inputs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Xinli, Ge, Wenhang, Lin, Jiantao, Feng, Jiawei, Xu, Lie, Zhao, HanFeng, Zhang, Shunsi, Chen, Ying-Cong
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910649347473408
author Xu, Xinli
Ge, Wenhang
Lin, Jiantao
Feng, Jiawei
Xu, Lie
Zhao, HanFeng
Zhang, Shunsi
Chen, Ying-Cong
author_facet Xu, Xinli
Ge, Wenhang
Lin, Jiantao
Feng, Jiawei
Xu, Lie
Zhao, HanFeng
Zhang, Shunsi
Chen, Ying-Cong
contents In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware text annotations. We utilize the strong reasoning capabilities of GPT-4V to generate 3D-aware text annotations. By analyzing four orthogonal views of an object arranged as tiled multi-view images, GPT-4V can produce text annotations that include 3D-aware information with spatial relationship. By integrating the control signal with proposed adaptive dual-control module, our model can generate multi-view images that correspond to the specified text. FlexGen supports multiple controllable capabilities, allowing users to modify text prompts to generate reasonable and corresponding unseen parts. Additionally, users can influence attributes such as appearance and material properties, including metallic and roughness. Extensive experiments demonstrate that our approach offers enhanced multiple controllability, marking a significant advancement over existing multi-view diffusion models. This work has substantial implications for fields requiring rapid and flexible 3D content creation, including game development, animation, and virtual reality. Project page: https://xxu068.github.io/flexgen.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10745
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FlexGen: Flexible Multi-View Generation from Text and Image Inputs
Xu, Xinli
Ge, Wenhang
Lin, Jiantao
Feng, Jiawei
Xu, Lie
Zhao, HanFeng
Zhang, Shunsi
Chen, Ying-Cong
Computer Vision and Pattern Recognition
Artificial Intelligence
In this work, we introduce FlexGen, a flexible framework designed to generate controllable and consistent multi-view images, conditioned on a single-view image, or a text prompt, or both. FlexGen tackles the challenges of controllable multi-view synthesis through additional conditioning on 3D-aware text annotations. We utilize the strong reasoning capabilities of GPT-4V to generate 3D-aware text annotations. By analyzing four orthogonal views of an object arranged as tiled multi-view images, GPT-4V can produce text annotations that include 3D-aware information with spatial relationship. By integrating the control signal with proposed adaptive dual-control module, our model can generate multi-view images that correspond to the specified text. FlexGen supports multiple controllable capabilities, allowing users to modify text prompts to generate reasonable and corresponding unseen parts. Additionally, users can influence attributes such as appearance and material properties, including metallic and roughness. Extensive experiments demonstrate that our approach offers enhanced multiple controllability, marking a significant advancement over existing multi-view diffusion models. This work has substantial implications for fields requiring rapid and flexible 3D content creation, including game development, animation, and virtual reality. Project page: https://xxu068.github.io/flexgen.github.io/.
title FlexGen: Flexible Multi-View Generation from Text and Image Inputs
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.10745