Structured 3D Latents for Scalable and Versatile 3D Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Jianfeng, Lv, Zelong, Xu, Sicheng, Deng, Yu, Wang, Ruicheng, Zhang, Bowen, Chen, Dong, Tong, Xin, Yang, Jiaolong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Native and Compact Structured Latents for 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2025)
by: Xiang, Jianfeng, et al.
Published: (2025)
Diffusion Models are Geometry Critics: Single Image 3D Editing Using Pre-Trained Diffusion Priors
by: Wang, Ruicheng, et al.
Published: (2024)
by: Wang, Ruicheng, et al.
Published: (2024)
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
by: Wang, Ruicheng, et al.
Published: (2025)
by: Wang, Ruicheng, et al.
Published: (2025)
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
by: Wang, Ruicheng, et al.
Published: (2024)
by: Wang, Ruicheng, et al.
Published: (2024)
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)
by: Liang, Huizhi, et al.
Published: (2026)
Map2World: Segment Map Conditioned Text to 3D World Generation
by: Chung, Jaeyoung, et al.
Published: (2026)
by: Chung, Jaeyoung, et al.
Published: (2026)
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image
by: Xu, Sicheng, et al.
Published: (2025)
by: Xu, Sicheng, et al.
Published: (2025)
ESGaussianFace: Emotional and Stylized Audio-Driven Facial Animation via 3D Gaussian Splatting
by: Ma, Chuhang, et al.
Published: (2026)
by: Ma, Chuhang, et al.
Published: (2026)
A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data
by: Xu, Yizhao, et al.
Published: (2026)
by: Xu, Yizhao, et al.
Published: (2026)
Compress3D: a Compressed Latent Space for 3D Generation from a Single Image
by: Zhang, Bowen, et al.
Published: (2024)
by: Zhang, Bowen, et al.
Published: (2024)
VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time
by: Xu, Sicheng, et al.
Published: (2024)
by: Xu, Sicheng, et al.
Published: (2024)
Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
by: Xu, Sicheng, et al.
Published: (2026)
by: Xu, Sicheng, et al.
Published: (2026)
LPA3D: 3D Room-Level Scene Generation from In-the-Wild Images
by: Yang, Ming-Jia, et al.
Published: (2025)
by: Yang, Ming-Jia, et al.
Published: (2025)
LN3DIFF++: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation
by: Lan, Yushi, et al.
Published: (2024)
by: Lan, Yushi, et al.
Published: (2024)
DC-Scene: Data-Centric Learning for 3D Scene Understanding
by: Huang, Ting, et al.
Published: (2025)
by: Huang, Ting, et al.
Published: (2025)
Brain3D: Generating 3D Objects from fMRI
by: Yang, Yuankun, et al.
Published: (2024)
by: Yang, Yuankun, et al.
Published: (2024)
InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
by: Xu, Sirui, et al.
Published: (2025)
by: Xu, Sirui, et al.
Published: (2025)
MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
by: Leng, Jiaqi, et al.
Published: (2026)
by: Leng, Jiaqi, et al.
Published: (2026)
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
by: Zhang, Ruicheng, et al.
Published: (2025)
by: Zhang, Ruicheng, et al.
Published: (2025)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
Versatile Video Tokenization with Generative 2D Gaussian Splatting
by: Chen, Zhenghao, et al.
Published: (2025)
by: Chen, Zhenghao, et al.
Published: (2025)
StructLDM: Structured Latent Diffusion for 3D Human Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion
by: Qiao, Yiran, et al.
Published: (2026)
by: Qiao, Yiran, et al.
Published: (2026)
PF3plat: Pose-Free Feed-Forward 3D Gaussian Splatting
by: Hong, Sunghwan, et al.
Published: (2024)
by: Hong, Sunghwan, et al.
Published: (2024)
Diverse 3D Human Pose Generation in Scenes based on Decoupled Structure
by: Dang, Bowen, et al.
Published: (2024)
by: Dang, Bowen, et al.
Published: (2024)
UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
by: He, Xufan, et al.
Published: (2025)
by: He, Xufan, et al.
Published: (2025)
SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
by: Xu, Mutian, et al.
Published: (2023)
by: Xu, Mutian, et al.
Published: (2023)
SS4D: Native 4D Generative Model via Structured Spacetime Latents
by: Li, Zhibing, et al.
Published: (2025)
by: Li, Zhibing, et al.
Published: (2025)
Omni-3DEdit: Generalized Versatile 3D Editing in One-Pass
by: Liyi, Chen, et al.
Published: (2026)
by: Liyi, Chen, et al.
Published: (2026)
ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
by: Chen, Honghua, et al.
Published: (2025)
by: Chen, Honghua, et al.
Published: (2025)
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
by: Liu, Shang, et al.
Published: (2025)
by: Liu, Shang, et al.
Published: (2025)
From One to More: Contextual Part Latents for 3D Generation
by: Dong, Shaocong, et al.
Published: (2025)
by: Dong, Shaocong, et al.
Published: (2025)
MeMix: Writing Less, Remembering More for Streaming 3D Reconstruction
by: Dong, Jiacheng, et al.
Published: (2026)
by: Dong, Jiacheng, et al.
Published: (2026)
DiGA3D: Coarse-to-Fine Diffusional Propagation of Geometry and Appearance for Versatile 3D Inpainting
by: Pan, Jingyi, et al.
Published: (2025)
by: Pan, Jingyi, et al.
Published: (2025)
Similar Items
-
Native and Compact Structured Latents for 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2025) -
Diffusion Models are Geometry Critics: Single Image 3D Editing Using Pre-Trained Diffusion Priors
by: Wang, Ruicheng, et al.
Published: (2024) -
MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
by: Wang, Ruicheng, et al.
Published: (2025) -
MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
by: Wang, Ruicheng, et al.
Published: (2024) -
HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
by: Liang, Huizhi, et al.
Published: (2026)