Video Perception Models for 3D Scene Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Rui, Zhai, Guangyao, Bauer, Zuria, Pollefeys, Marc, Tombari, Federico, Guibas, Leonidas, Huang, Gao, Engelmann, Francis |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SuperDec: 3D Scene Decomposition with Superquadric Primitives
by: Fedele, Elisabetta, et al.
Published: (2025)
by: Fedele, Elisabetta, et al.
Published: (2025)
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
by: Fedele, Elisabetta, et al.
Published: (2025)
by: Fedele, Elisabetta, et al.
Published: (2025)
SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs
by: Miao, Yang, et al.
Published: (2024)
by: Miao, Yang, et al.
Published: (2024)
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
Robust Human Registration with Body Part Segmentation on Noisy Point Clouds
by: Lascheit, Kai, et al.
Published: (2025)
by: Lascheit, Kai, et al.
Published: (2025)
Lost & Found: Tracking Changes from Egocentric Observations in 3D Dynamic Scene Graphs
by: Behrens, Tjark, et al.
Published: (2024)
by: Behrens, Tjark, et al.
Published: (2024)
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
by: Delitzas, Alexandros, et al.
Published: (2026)
by: Delitzas, Alexandros, et al.
Published: (2026)
P2P-Bridge: Diffusion Bridges for 3D Point Cloud Denoising
by: Vogel, Mathias, et al.
Published: (2024)
by: Vogel, Mathias, et al.
Published: (2024)
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
by: Maillard, Léopold, et al.
Published: (2026)
by: Maillard, Léopold, et al.
Published: (2026)
FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning
by: Fu, Zhengyu, et al.
Published: (2026)
by: Fu, Zhengyu, et al.
Published: (2026)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
by: Shang, Haoliang, et al.
Published: (2025)
by: Shang, Haoliang, et al.
Published: (2025)
Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos
by: Stearns, Colton, et al.
Published: (2024)
by: Stearns, Colton, et al.
Published: (2024)
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
Spot-Compose: A Framework for Open-Vocabulary Object Retrieval and Drawer Manipulation in Point Clouds
by: Lemke, Oliver, et al.
Published: (2024)
by: Lemke, Oliver, et al.
Published: (2024)
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
by: Yang, Yung-Hsu, et al.
Published: (2025)
by: Yang, Yung-Hsu, et al.
Published: (2025)
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024)
by: Huang, Ian, et al.
Published: (2024)
Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
by: Di Lorenzo, Gaia, et al.
Published: (2025)
by: Di Lorenzo, Gaia, et al.
Published: (2025)
UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections
by: Wang, Fangjinhua, et al.
Published: (2023)
by: Wang, Fangjinhua, et al.
Published: (2023)
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
by: Büchner, Martin, et al.
Published: (2026)
by: Büchner, Martin, et al.
Published: (2026)
MaRINeR: Enhancing Novel Views by Matching Rendered Images with Nearby References
by: Bösiger, Lukas, et al.
Published: (2024)
by: Bösiger, Lukas, et al.
Published: (2024)
ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding
by: Ji, Guangda, et al.
Published: (2024)
by: Ji, Guangda, et al.
Published: (2024)
HouseLayout3D: A Benchmark and Training-Free Baseline for 3D Layout Estimation in the Wild
by: Bieri, Valentin, et al.
Published: (2025)
by: Bieri, Valentin, et al.
Published: (2025)
Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor Spaces
by: Zhang, Chenyangguang, et al.
Published: (2025)
by: Zhang, Chenyangguang, et al.
Published: (2025)
GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering
by: Li, Yanyan, et al.
Published: (2024)
by: Li, Yanyan, et al.
Published: (2024)
SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation
by: Siegel, Peter, et al.
Published: (2025)
by: Siegel, Peter, et al.
Published: (2025)
OpenDAS: Open-Vocabulary Domain Adaptation for 2D and 3D Segmentation
by: Yilmaz, Gonca, et al.
Published: (2024)
by: Yilmaz, Gonca, et al.
Published: (2024)
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
by: Hu, Xinggang, et al.
Published: (2026)
by: Hu, Xinggang, et al.
Published: (2026)
SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs
by: Zhai, Guangyao, et al.
Published: (2023)
by: Zhai, Guangyao, et al.
Published: (2023)
Search3D: Hierarchical Open-Vocabulary 3D Segmentation
by: Takmaz, Ayca, et al.
Published: (2024)
by: Takmaz, Ayca, et al.
Published: (2024)
Mixed Diffusion for 3D Indoor Scene Synthesis
by: Hu, Siyi, et al.
Published: (2024)
by: Hu, Siyi, et al.
Published: (2024)
REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
by: Huang, Zhao, et al.
Published: (2025)
by: Huang, Zhao, et al.
Published: (2025)
OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
by: Deng, Zilong, et al.
Published: (2026)
by: Deng, Zilong, et al.
Published: (2026)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
by: Meier, Johannes, et al.
Published: (2025)
by: Meier, Johannes, et al.
Published: (2025)
OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
Gaussians-to-Life: Text-Driven Animation of 3D Gaussian Splatting Scenes
by: Wimmer, Thomas, et al.
Published: (2024)
by: Wimmer, Thomas, et al.
Published: (2024)
OCH3R: Object-Centric Holistic 3D Reconstruction
by: Du, Yi, et al.
Published: (2026)
by: Du, Yi, et al.
Published: (2026)
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2024)
by: Zhai, Guangyao, et al.
Published: (2024)
UnLoc: Leveraging Depth Uncertainties for Floorplan Localization
by: Wüest, Matthias, et al.
Published: (2025)
by: Wüest, Matthias, et al.
Published: (2025)
CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization
by: Ackermann, Jan, et al.
Published: (2025)
by: Ackermann, Jan, et al.
Published: (2025)
Similar Items
-
SuperDec: 3D Scene Decomposition with Superquadric Primitives
by: Fedele, Elisabetta, et al.
Published: (2025) -
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
by: Fedele, Elisabetta, et al.
Published: (2025) -
SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs
by: Miao, Yang, et al.
Published: (2024) -
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views
by: Engelmann, Francis, et al.
Published: (2024) -
Robust Human Registration with Body Part Segmentation on Noisy Point Clouds
by: Lascheit, Kai, et al.
Published: (2025)