Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Burgess, James, Wang, Kuan-Chieh, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
by: Fang, Shaoheng, et al.
Published: (2025)
by: Fang, Shaoheng, et al.
Published: (2025)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
by: Lozano, Alejandro, et al.
Published: (2024)
by: Lozano, Alejandro, et al.
Published: (2024)
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
BetterScene: 3D Scene Synthesis with Representation-Aligned Generative Model
by: Han, Yuci, et al.
Published: (2026)
by: Han, Yuci, et al.
Published: (2026)
MagicDrive: Street View Generation with Diverse 3D Geometry Control
by: Gao, Ruiyuan, et al.
Published: (2023)
by: Gao, Ruiyuan, et al.
Published: (2023)
3D Diffuser Actor: Policy Diffusion with 3D Scene Representations
by: Ke, Tsung-Wei, et al.
Published: (2024)
by: Ke, Tsung-Wei, et al.
Published: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
Quantum Implicit Neural Representations for 3D Scene Reconstruction and Novel View Synthesis
by: Cordero, Yeray, et al.
Published: (2025)
by: Cordero, Yeray, et al.
Published: (2025)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
2D Representation for Unguided Single-View 3D Super-Resolution in Real-Time
by: Mas, Ignasi, et al.
Published: (2025)
by: Mas, Ignasi, et al.
Published: (2025)
Shape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
by: Feng, Tuo, et al.
Published: (2024)
by: Feng, Tuo, et al.
Published: (2024)
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024)
by: Rau, Anita, et al.
Published: (2024)
Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model
by: Zhang, Shengjun, et al.
Published: (2025)
by: Zhang, Shengjun, et al.
Published: (2025)
HexPlane Representation for 3D Semantic Scene Understanding
by: Chen, Zeren, et al.
Published: (2025)
by: Chen, Zeren, et al.
Published: (2025)
Enhancing Monocular 3D Scene Completion with Diffusion Model
by: Song, Changlin, et al.
Published: (2025)
by: Song, Changlin, et al.
Published: (2025)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
by: Lee, Dong In, et al.
Published: (2024)
by: Lee, Dong In, et al.
Published: (2024)
LT3SD: Latent Trees for 3D Scene Diffusion
by: Meng, Quan, et al.
Published: (2024)
by: Meng, Quan, et al.
Published: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
by: Mou, Linzhan, et al.
Published: (2024)
by: Mou, Linzhan, et al.
Published: (2024)
3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion Priors
by: Liu, Xi, et al.
Published: (2024)
by: Liu, Xi, et al.
Published: (2024)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024)
by: Li, Xiaolong, et al.
Published: (2024)
Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion
by: Liang, Li, et al.
Published: (2025)
by: Liang, Li, et al.
Published: (2025)
3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation
by: Chen, Hansheng, et al.
Published: (2024)
by: Chen, Hansheng, et al.
Published: (2024)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
by: Zhang, Frank, et al.
Published: (2024)
by: Zhang, Frank, et al.
Published: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
PolarBEVDet: Exploring Polar Representation for Multi-View 3D Object Detection in Bird's-Eye-View
by: Yu, Zichen, et al.
Published: (2024)
by: Yu, Zichen, et al.
Published: (2024)
RoomCraft: Controllable and Complete 3D Indoor Scene Generation
by: Zhou, Mengqi, et al.
Published: (2025)
by: Zhou, Mengqi, et al.
Published: (2025)
Pomo3D: 3D-Aware Portrait Accessorizing and More
by: Liu, Tzu-Chieh, et al.
Published: (2024)
by: Liu, Tzu-Chieh, et al.
Published: (2024)
Graph Canvas for Controllable 3D Scene Generation
by: Liu, Libin, et al.
Published: (2024)
by: Liu, Libin, et al.
Published: (2024)
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
by: Sun, Wenqiang, et al.
Published: (2024)
by: Sun, Wenqiang, et al.
Published: (2024)
Template-Free Single-View 3D Human Digitalization with Diffusion-Guided LRM
by: Weng, Zhenzhen, et al.
Published: (2024)
by: Weng, Zhenzhen, et al.
Published: (2024)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
Magic-Boost: Boost 3D Generation with Multi-View Conditioned Diffusion
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
by: Zheng, Haowen, et al.
Published: (2024)
by: Zheng, Haowen, et al.
Published: (2024)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
Similar Items
-
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026) -
Motion Diffusion-Guided 3D Global HMR from a Dynamic Camera
by: Heo, Jaewoo, et al.
Published: (2024) -
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
by: Bravo-Sánchez, Laura, et al.
Published: (2024) -
MVRoom: Controllable 3D Indoor Scene Generation with Multi-View Diffusion Models
by: Fang, Shaoheng, et al.
Published: (2025) -
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
by: Lozano, Alejandro, et al.
Published: (2024)