DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Qihao, Zhang, Yi, Bai, Song, Kortylewski, Adam, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
Generating Images with 3D Annotations Using Diffusion Models
by: Ma, Wufei, et al.
Published: (2023)
by: Ma, Wufei, et al.
Published: (2023)
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
by: Shi, Jiaxin, et al.
Published: (2026)
by: Shi, Jiaxin, et al.
Published: (2026)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
by: Jesslen, Artur, et al.
Published: (2026)
by: Jesslen, Artur, et al.
Published: (2026)
Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Space
by: Sommer, Leonhard, et al.
Published: (2025)
by: Sommer, Leonhard, et al.
Published: (2025)
A Bayesian Approach to OOD Robustness in Image Classification
by: Kaushik, Prakhar, et al.
Published: (2024)
by: Kaushik, Prakhar, et al.
Published: (2024)
Animal3D: A Comprehensive Dataset of 3D Animal Pose and Shape
by: Xu, Jiacong, et al.
Published: (2023)
by: Xu, Jiacong, et al.
Published: (2023)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
by: Zhong, Shanshan, et al.
Published: (2025)
by: Zhong, Shanshan, et al.
Published: (2025)
DatasetNeRF: Efficient 3D-aware Data Factory with Generative Radiance Fields
by: Chi, Yu, et al.
Published: (2023)
by: Chi, Yu, et al.
Published: (2023)
Prompt-Based Exemplar Super-Compression and Regeneration for Class-Incremental Learning
by: Duan, Ruxiao, et al.
Published: (2023)
by: Duan, Ruxiao, et al.
Published: (2023)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
by: Paul, Soumava, et al.
Published: (2026)
by: Paul, Soumava, et al.
Published: (2026)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
by: Guo, Weijie, et al.
Published: (2025)
by: Guo, Weijie, et al.
Published: (2025)
FaceGPT: Self-supervised Learning to Chat about 3D Human Faces
by: Wang, Haoran, et al.
Published: (2024)
by: Wang, Haoran, et al.
Published: (2024)
Unsupervised Learning of Category-Level 3D Pose from Object-Centric Videos
by: Sommer, Leonhard, et al.
Published: (2024)
by: Sommer, Leonhard, et al.
Published: (2024)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
by: Sheung, Eddie Pokming, et al.
Published: (2025)
by: Sheung, Eddie Pokming, et al.
Published: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
by: Huang, Tianyu, et al.
Published: (2023)
by: Huang, Tianyu, et al.
Published: (2023)
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
by: Sommer, Leonhard, et al.
Published: (2026)
by: Sommer, Leonhard, et al.
Published: (2026)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
by: Kaushik, Prakhar, et al.
Published: (2024)
by: Kaushik, Prakhar, et al.
Published: (2024)
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
by: Pham, Nhi, et al.
Published: (2025)
by: Pham, Nhi, et al.
Published: (2025)
Every9D-21M: Large-Scale Real-World 9D Canonicalization of Everyday Objects
by: Sommer, Leonhard, et al.
Published: (2026)
by: Sommer, Leonhard, et al.
Published: (2026)
Name That Part: 3D Part Segmentation and Naming
by: Paul, Soumava, et al.
Published: (2025)
by: Paul, Soumava, et al.
Published: (2025)
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
by: Chen, Qi, et al.
Published: (2026)
by: Chen, Qi, et al.
Published: (2026)
PI3D: Efficient Text-to-3D Generation with Pseudo-Image Diffusion
by: Liu, Ying-Tian, et al.
Published: (2023)
by: Liu, Ying-Tian, et al.
Published: (2023)
VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
by: Wang, Angtian, et al.
Published: (2022)
by: Wang, Angtian, et al.
Published: (2022)
NOVUM: Neural Object Volumes for Robust Object Classification
by: Jesslen, Artur, et al.
Published: (2023)
by: Jesslen, Artur, et al.
Published: (2023)
Learning a Category-level Object Pose Estimator without Pose Annotations
by: Tian, Fengrui, et al.
Published: (2024)
by: Tian, Fengrui, et al.
Published: (2024)
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
by: Guo, Zhanqiang, et al.
Published: (2024)
by: Guo, Zhanqiang, et al.
Published: (2024)
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
by: Fan, Chenyang, et al.
Published: (2026)
by: Fan, Chenyang, et al.
Published: (2026)
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
Turbo3D: Ultra-fast Text-to-3D Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
PAT3D: Physics-Augmented Text-to-3D Scene Generation
by: Lin, Guying, et al.
Published: (2025)
by: Lin, Guying, et al.
Published: (2025)
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
Dive3D: Diverse Distillation-based Text-to-3D Generation via Score Implicit Matching
by: Bai, Weimin, et al.
Published: (2025)
by: Bai, Weimin, et al.
Published: (2025)
Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D Prior
by: Chen, Cheng, et al.
Published: (2024)
by: Chen, Cheng, et al.
Published: (2024)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
by: Wu, Xiuzhe, et al.
Published: (2024)
by: Wu, Xiuzhe, et al.
Published: (2024)
Similar Items
-
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
by: Ma, Wufei, et al.
Published: (2024) -
Generating Images with 3D Annotations Using Diffusion Models
by: Ma, Wufei, et al.
Published: (2023) -
PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views
by: Shi, Jiaxin, et al.
Published: (2026) -
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
by: Liu, Qihao, et al.
Published: (2025) -
Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
by: Jesslen, Artur, et al.
Published: (2026)