Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Zhenggang, Wang, Yuehao, Fan, Yuchen, Chen, Jun-Kun, Yeh, Yu-Ying, Sohn, Kihyuk, Wang, Zhangyang, Huang, Qixing, Schwing, Alexander, Ranjan, Rakesh, Wang, Dilin, Yan, Zhicheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
by: Tang, Zhenggang, et al.
Published: (2024)
by: Tang, Zhenggang, et al.
Published: (2024)
Steepest Descent Density Control for Compact 3D Gaussian Splatting
by: Wang, Peihao, et al.
Published: (2025)
by: Wang, Peihao, et al.
Published: (2025)
MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction
by: Tang, Shitao, et al.
Published: (2024)
by: Tang, Shitao, et al.
Published: (2024)
Taming Mode Collapse in Score Distillation for Text-to-3D Generation
by: Wang, Peihao, et al.
Published: (2023)
by: Wang, Peihao, et al.
Published: (2023)
SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein Identity
by: Wang, Peihao, et al.
Published: (2023)
by: Wang, Peihao, et al.
Published: (2023)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion Models
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
AssetGen: Deployable 3D Asset Generation at Interactive Speed
by: Wang, Dilin, et al.
Published: (2026)
by: Wang, Dilin, et al.
Published: (2026)
DreamFlow: High-Quality Text-to-3D Generation by Approximating Probability Flow
by: Lee, Kyungmin, et al.
Published: (2024)
by: Lee, Kyungmin, et al.
Published: (2024)
3D Mesh Editing using Masked LRMs
by: Gao, Will, et al.
Published: (2024)
by: Gao, Will, et al.
Published: (2024)
WorldGen: From Text to Traversable and Interactive 3D Worlds
by: Wang, Dilin, et al.
Published: (2025)
by: Wang, Dilin, et al.
Published: (2025)
Zero-Painter: Training-Free Layout Control for Text-to-Image Synthesis
by: Ohanyan, Marianna, et al.
Published: (2024)
by: Ohanyan, Marianna, et al.
Published: (2024)
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
by: Tang, Zhenggang, et al.
Published: (2024)
by: Tang, Zhenggang, et al.
Published: (2024)
Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
by: Xu, Xingqian, et al.
Published: (2022)
by: Xu, Xingqian, et al.
Published: (2022)
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025)
by: Zheng, Haozhen, et al.
Published: (2025)
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
by: Zheng, Zirui, et al.
Published: (2025)
by: Zheng, Zirui, et al.
Published: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
AutoPartGen: Autogressive 3D Part Generation and Discovery
by: Chen, Minghao, et al.
Published: (2025)
by: Chen, Minghao, et al.
Published: (2025)
MoVideo: Motion-Aware Video Generation with Diffusion Models
by: Liang, Jingyun, et al.
Published: (2023)
by: Liang, Jingyun, et al.
Published: (2023)
A Mixed Integer‐Linear Programming Model for Solving the Hydroelectric Unit Maintenance Scheduling Problem
by: Yuehao Tang
Published: (2025)
by: Yuehao Tang
Published: (2025)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
by: Liu, Zhijun, et al.
Published: (2024)
by: Liu, Zhijun, et al.
Published: (2024)
NeRFDeformer: NeRF Transformation from a Single View via 3D Scene Flows
by: Tang, Zhenggang, et al.
Published: (2024)
by: Tang, Zhenggang, et al.
Published: (2024)
MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
Oscillation Inversion: Understand the structure of Large Flow Model through the Lens of Inversion Method
by: Zheng, Yan, et al.
Published: (2024)
by: Zheng, Yan, et al.
Published: (2024)
Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
by: Wang, Peihao, et al.
Published: (2024)
by: Wang, Peihao, et al.
Published: (2024)
Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction
by: Saito, Kuniaki, et al.
Published: (2024)
by: Saito, Kuniaki, et al.
Published: (2024)
Robust Disaster Assessment from Aerial Imagery Using Text-to-Image Synthetic Data
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation
by: Zheng, Guangcong, et al.
Published: (2023)
by: Zheng, Guangcong, et al.
Published: (2023)
Make-A-Texture: Fast Shape-Aware Texture Generation in 3 Seconds
by: Xiang, Xiaoyu, et al.
Published: (2024)
by: Xiang, Xiaoyu, et al.
Published: (2024)
STAY Diffusion: Styled Layout Diffusion Model for Diverse Layout-to-Image Generation
by: Wang, Ruyu, et al.
Published: (2025)
by: Wang, Ruyu, et al.
Published: (2025)
SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
by: Wang, Sen, et al.
Published: (2025)
by: Wang, Sen, et al.
Published: (2025)
Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models
by: Lee, Kihyuk
Published: (2026)
by: Lee, Kihyuk
Published: (2026)
Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model
by: Lee, Kihyuk
Published: (2026)
by: Lee, Kihyuk
Published: (2026)
On Inductive Biases That Enable Generalization of Diffusion Transformers
by: An, Jie, et al.
Published: (2024)
by: An, Jie, et al.
Published: (2024)
LiteGE: Lightweight Geodesic Embedding for Efficient Geodesics Computation and Non-Isometric Shape Correspondence
by: Adikusuma, Yohanes Yudhi, et al.
Published: (2025)
by: Adikusuma, Yohanes Yudhi, et al.
Published: (2025)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
Recurrent Diffusion for Large-Scale Parameter Generation
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
LoCoCo: Dropping In Convolutions for Long Context Compression
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
Variational Masked Diffusion Models
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
FlexGS: Train Once, Deploy Everywhere with Many-in-One Flexible 3D Gaussian Splatting
by: Liu, Hengyu, et al.
Published: (2025)
by: Liu, Hengyu, et al.
Published: (2025)
Similar Items
-
MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
by: Tang, Zhenggang, et al.
Published: (2024) -
Steepest Descent Density Control for Compact 3D Gaussian Splatting
by: Wang, Peihao, et al.
Published: (2025) -
MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction
by: Tang, Shitao, et al.
Published: (2024) -
Taming Mode Collapse in Score Distillation for Text-to-3D Generation
by: Wang, Peihao, et al.
Published: (2023) -
SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein Identity
by: Wang, Peihao, et al.
Published: (2023)