Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Meng, Xuyi, Wang, Chen, Lei, Jiahui, Daniilidis, Kostas, Gu, Jiatao, Liu, Lingjie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
by: Mou, Linzhan, et al.
Published: (2025)
by: Mou, Linzhan, et al.
Published: (2025)
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
by: Song, Yunzhou, et al.
Published: (2024)
by: Song, Yunzhou, et al.
Published: (2024)
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
by: Wang, Yufu, et al.
Published: (2024)
by: Wang, Yufu, et al.
Published: (2024)
DynMF: Neural Motion Factorization for Real-time Dynamic View Synthesis with 3D Gaussian Splatting
by: Kratimenos, Agelos, et al.
Published: (2023)
by: Kratimenos, Agelos, et al.
Published: (2023)
Track Everything Everywhere Fast and Robustly
by: Song, Yunzhou, et al.
Published: (2024)
by: Song, Yunzhou, et al.
Published: (2024)
GECO: Generative Image-to-3D within a SECOnd
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
by: Lin, Dixuan, et al.
Published: (2025)
by: Lin, Dixuan, et al.
Published: (2025)
MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
by: Lei, Jiahui, et al.
Published: (2024)
by: Lei, Jiahui, et al.
Published: (2024)
Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
by: Li, Yiqian, et al.
Published: (2025)
by: Li, Yiqian, et al.
Published: (2025)
FisherRF: Active View Selection and Uncertainty Quantification for Radiance Fields using Fisher Information
by: Jiang, Wen, et al.
Published: (2023)
by: Jiang, Wen, et al.
Published: (2023)
Bring the Power of Diffusion Model to Defect Detection
by: Yu, Xuyi
Published: (2024)
by: Yu, Xuyi
Published: (2024)
Electrostatics-Inspired Surface Reconstruction (EISR): Recovering 3D Shapes as a Superposition of Poisson's PDE Solutions
by: Patiño, Diego, et al.
Published: (2026)
by: Patiño, Diego, et al.
Published: (2026)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
by: Zhang, Ruijun, et al.
Published: (2026)
by: Zhang, Ruijun, et al.
Published: (2026)
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting
by: Strong, Matthew, et al.
Published: (2024)
by: Strong, Matthew, et al.
Published: (2024)
World-consistent Video Diffusion with Explicit 3D Modeling
by: Zhang, Qihang, et al.
Published: (2024)
by: Zhang, Qihang, et al.
Published: (2024)
LN3DIFF++: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation
by: Lan, Yushi, et al.
Published: (2024)
by: Lan, Yushi, et al.
Published: (2024)
PhysHMR: Learning Humanoid Control Policies from Vision for Physically Plausible Human Motion Reconstruction
by: Feng, Qiao, et al.
Published: (2025)
by: Feng, Qiao, et al.
Published: (2025)
Un-EVIMO: Unsupervised Event-Based Independent Motion Segmentation
by: Wang, Ziyun, et al.
Published: (2023)
by: Wang, Ziyun, et al.
Published: (2023)
Multimodal LLM Guided Exploration and Active Mapping using Fisher Information
by: Jiang, Wen, et al.
Published: (2024)
by: Jiang, Wen, et al.
Published: (2024)
Pantheon360: Taming Digital Twin Generation via 3D-Aware 360° Video Diffusion
by: Chen, Ting-Hsuan, et al.
Published: (2026)
by: Chen, Ting-Hsuan, et al.
Published: (2026)
EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decomposition
by: Hu, Yihan, et al.
Published: (2025)
by: Hu, Yihan, et al.
Published: (2025)
Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
by: Song, Chenxi, et al.
Published: (2025)
by: Song, Chenxi, et al.
Published: (2025)
BiEquiFormer: Bi-Equivariant Representations for Global Point Cloud Registration
by: Pertigkiozoglou, Stefanos, et al.
Published: (2024)
by: Pertigkiozoglou, Stefanos, et al.
Published: (2024)
HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
by: Zhou, Haiyang, et al.
Published: (2025)
by: Zhou, Haiyang, et al.
Published: (2025)
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models
by: Miao, Kevin, et al.
Published: (2024)
by: Miao, Kevin, et al.
Published: (2024)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
by: Bahmani, Sherwin, et al.
Published: (2024)
by: Bahmani, Sherwin, et al.
Published: (2024)
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
by: Das, Richeek, et al.
Published: (2025)
by: Das, Richeek, et al.
Published: (2025)
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
by: Liu, Zixian, et al.
Published: (2026)
by: Liu, Zixian, et al.
Published: (2026)
Leveraging Pretrained Diffusion Models for Zero-Shot Part Assembly
by: Zhang, Ruiyuan, et al.
Published: (2025)
by: Zhang, Ruiyuan, et al.
Published: (2025)
RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation
by: Huang, Hanzhuo, et al.
Published: (2026)
by: Huang, Hanzhuo, et al.
Published: (2026)
Many-to-many Image Generation with Auto-regressive Diffusion Models
by: Shen, Ying, et al.
Published: (2024)
by: Shen, Ying, et al.
Published: (2024)
V3D: Video Diffusion Models are Effective 3D Generators
by: Chen, Zilong, et al.
Published: (2024)
by: Chen, Zilong, et al.
Published: (2024)
Lift3D Foundation Policy: Lifting 2D Large-Scale Pretrained Models for Robust 3D Robotic Manipulation
by: Jia, Yueru, et al.
Published: (2024)
by: Jia, Yueru, et al.
Published: (2024)
$SE(3)$ Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation
by: Xu, Yinshuang, et al.
Published: (2024)
by: Xu, Yinshuang, et al.
Published: (2024)
Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer
by: Bui, Minh, et al.
Published: (2024)
by: Bui, Minh, et al.
Published: (2024)
VGGT-HPE: Reframing Head Pose Estimation as Relative Pose Prediction
by: Vasileiou, Vasiliki, et al.
Published: (2026)
by: Vasileiou, Vasiliki, et al.
Published: (2026)
RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion
by: Shriram, Jaidev, et al.
Published: (2024)
by: Shriram, Jaidev, et al.
Published: (2024)
Similar Items
-
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
by: Mou, Linzhan, et al.
Published: (2025) -
StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation
by: Li, Haodong, et al.
Published: (2025) -
HDGS: Textured 2D Gaussian Splatting for Enhanced Scene Rendering
by: Song, Yunzhou, et al.
Published: (2024) -
TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
by: Wang, Yufu, et al.
Published: (2024) -
DynMF: Neural Motion Factorization for Real-time Dynamic View Synthesis with 3D Gaussian Splatting
by: Kratimenos, Agelos, et al.
Published: (2023)