Diffusion Models in 3D Vision: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhen, Li, Dongyuan, Wu, Yaozu, He, Tianyu, Bian, Jiang, Jiang, Renhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
von: Wu, Haoyu, et al.
Veröffentlicht: (2025)
von: Wu, Haoyu, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Inverse Rewards for World Model Post-training
von: Ye, Yang, et al.
Veröffentlicht: (2025)
von: Ye, Yang, et al.
Veröffentlicht: (2025)
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
Compositional 3D-aware Video Generation with LLM Director
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024)
GeoGS3D: Single-view 3D Reconstruction via Geometric-aware Diffusion Model and Gaussian Splatting
von: Feng, Qijun, et al.
Veröffentlicht: (2024)
von: Feng, Qijun, et al.
Veröffentlicht: (2024)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
von: Ma, Weijian, et al.
Veröffentlicht: (2026)
von: Ma, Weijian, et al.
Veröffentlicht: (2026)
AR4D: Autoregressive 4D Generation from Monocular Videos
von: Zhu, Hanxin, et al.
Veröffentlicht: (2025)
von: Zhu, Hanxin, et al.
Veröffentlicht: (2025)
End-to-End Rate-Distortion Optimized 3D Gaussian Representation
von: Wang, Henan, et al.
Veröffentlicht: (2024)
von: Wang, Henan, et al.
Veröffentlicht: (2024)
Playing with Transformer at 30+ FPS via Next-Frame Diffusion
von: Cheng, Xinle, et al.
Veröffentlicht: (2025)
von: Cheng, Xinle, et al.
Veröffentlicht: (2025)
LIVE: Long-horizon Interactive Video World Modeling
von: Huang, Junchao, et al.
Veröffentlicht: (2026)
von: Huang, Junchao, et al.
Veröffentlicht: (2026)
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
von: Guo, Junliang, et al.
Veröffentlicht: (2025)
von: Guo, Junliang, et al.
Veröffentlicht: (2025)
A Survey on Vision Autoregressive Model
von: Jiang, Kai, et al.
Veröffentlicht: (2024)
von: Jiang, Kai, et al.
Veröffentlicht: (2024)
CatFree3D: Category-agnostic 3D Object Detection with Diffusion
von: Bian, Wenjing, et al.
Veröffentlicht: (2024)
von: Bian, Wenjing, et al.
Veröffentlicht: (2024)
A Survey on Personalized Content Synthesis with Diffusion Models
von: Zhang, Xulu, et al.
Veröffentlicht: (2024)
von: Zhang, Xulu, et al.
Veröffentlicht: (2024)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
von: Zhang, Wentao, et al.
Veröffentlicht: (2024)
von: Zhang, Wentao, et al.
Veröffentlicht: (2024)
TiC: Exploring Vision Transformer in Convolution
von: Zhang, Song, et al.
Veröffentlicht: (2023)
von: Zhang, Song, et al.
Veröffentlicht: (2023)
GaussianSR: 3D Gaussian Super-Resolution with 2D Diffusion Priors
von: Yu, Xiqian, et al.
Veröffentlicht: (2024)
von: Yu, Xiqian, et al.
Veröffentlicht: (2024)
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
von: Xu, Wan, et al.
Veröffentlicht: (2023)
von: Xu, Wan, et al.
Veröffentlicht: (2023)
A Survey On Text-to-3D Contents Generation In The Wild
von: Jiang, Chenhan
Veröffentlicht: (2024)
von: Jiang, Chenhan
Veröffentlicht: (2024)
Diffusion Models in Low-Level Vision: A Survey
von: He, Chunming, et al.
Veröffentlicht: (2024)
von: He, Chunming, et al.
Veröffentlicht: (2024)
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
von: Bai, Jianhong, et al.
Veröffentlicht: (2024)
von: Bai, Jianhong, et al.
Veröffentlicht: (2024)
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
von: Huang, Tianyu, et al.
Veröffentlicht: (2025)
von: Huang, Tianyu, et al.
Veröffentlicht: (2025)
SceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RL
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
von: Zhao, Yang, et al.
Veröffentlicht: (2026)
Rethinking Preference Alignment for Diffusion Models with Classifier-Free Guidance
von: Jiang, Zhou, et al.
Veröffentlicht: (2026)
von: Jiang, Zhou, et al.
Veröffentlicht: (2026)
A Survey on Video Diffusion Models
von: Xing, Zhen, et al.
Veröffentlicht: (2023)
von: Xing, Zhen, et al.
Veröffentlicht: (2023)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Beyond Pixel Histories: World Models with Persistent 3D State
von: Garcin, Samuel, et al.
Veröffentlicht: (2026)
von: Garcin, Samuel, et al.
Veröffentlicht: (2026)
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion
von: Jiang, Yanqin, et al.
Veröffentlicht: (2024)
von: Jiang, Yanqin, et al.
Veröffentlicht: (2024)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
von: Weng, Zejia, et al.
Veröffentlicht: (2024)
von: Weng, Zejia, et al.
Veröffentlicht: (2024)
Comprehensive Survey of Model Compression and Speed up for Vision Transformers
von: Chen, Feiyang, et al.
Veröffentlicht: (2024)
von: Chen, Feiyang, et al.
Veröffentlicht: (2024)
Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis
von: Lai, Haoran, et al.
Veröffentlicht: (2026)
von: Lai, Haoran, et al.
Veröffentlicht: (2026)
PQD: Post-training Quantization for Efficient Diffusion Models
von: Ye, Jiaojiao, et al.
Veröffentlicht: (2024)
von: Ye, Jiaojiao, et al.
Veröffentlicht: (2024)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
Hierarchical Masked 3D Diffusion Model for Video Outpainting
von: Fan, Fanda, et al.
Veröffentlicht: (2023)
von: Fan, Fanda, et al.
Veröffentlicht: (2023)
BGDB: Bernoulli-Gaussian Decision Block with Improved Denoising Diffusion Probabilistic Models
von: Sun, Chengkun, et al.
Veröffentlicht: (2024)
von: Sun, Chengkun, et al.
Veröffentlicht: (2024)
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
von: Jiang, Dengyang, et al.
Veröffentlicht: (2026)
von: Jiang, Dengyang, et al.
Veröffentlicht: (2026)
VLM-3D:End-to-End Vision-Language Models for Open-World 3D Perception
von: Chang, Fuhao, et al.
Veröffentlicht: (2025)
von: Chang, Fuhao, et al.
Veröffentlicht: (2025)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
von: Nguyen, Hieu T., et al.
Veröffentlicht: (2024)
von: Nguyen, Hieu T., et al.
Veröffentlicht: (2024)
Can Vision Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective
von: An, Arctanx, et al.
Veröffentlicht: (2026)
von: An, Arctanx, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
von: Wu, Haoyu, et al.
Veröffentlicht: (2025) -
Reinforcement Learning with Inverse Rewards for World Model Post-training
von: Ye, Yang, et al.
Veröffentlicht: (2025) -
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
von: Deng, Jiajun, et al.
Veröffentlicht: (2025) -
Compositional 3D-aware Video Generation with LLM Director
von: Zhu, Hanxin, et al.
Veröffentlicht: (2024) -
GeoGS3D: Single-view 3D Reconstruction via Geometric-aware Diffusion Model and Gaussian Splatting
von: Feng, Qijun, et al.
Veröffentlicht: (2024)