VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
Fuente:
arXiv
Saved in:
| Main Authors: | Kuang, Zhengfei, Lin, Rui, Zhao, Long, Wetzstein, Gordon, Xie, Saining, Woo, Sanghyun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
by: Chen, Jiacheng, et al.
Published: (2025)
by: Chen, Jiacheng, et al.
Published: (2025)
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
by: Ackermann, Jan, et al.
Published: (2026)
by: Ackermann, Jan, et al.
Published: (2026)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
by: Yenphraphai, Jiraphon, et al.
Published: (2024)
Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark
by: Kuang, Zhengfei, et al.
Published: (2023)
by: Kuang, Zhengfei, et al.
Published: (2023)
EVT: Efficient View Transformation for Multi-Modal 3D Object Detection
by: Lee, Yongjin, et al.
Published: (2024)
by: Lee, Yongjin, et al.
Published: (2024)
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
MultiDepth: Multi-Sample Priors for Refining Monocular Metric Depth Estimations in Indoor Scenes
by: Byun, Sanghyun, et al.
Published: (2024)
by: Byun, Sanghyun, et al.
Published: (2024)
Multi-Modal Decouple and Recouple Network for Robust 3D Object Detection
by: Ding, Rui, et al.
Published: (2026)
by: Ding, Rui, et al.
Published: (2026)
RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object Detection
by: Ding, Rui, et al.
Published: (2026)
by: Ding, Rui, et al.
Published: (2026)
3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models
by: Zhang, Yuhan, et al.
Published: (2025)
by: Zhang, Yuhan, et al.
Published: (2025)
CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
by: Fang, Ye, et al.
Published: (2024)
by: Fang, Ye, et al.
Published: (2024)
LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation
by: Yang, Shuai, et al.
Published: (2024)
by: Yang, Shuai, et al.
Published: (2024)
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
by: Gu, Yuming, et al.
Published: (2025)
by: Gu, Yuming, et al.
Published: (2025)
GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
by: Zhong, Yao, et al.
Published: (2025)
by: Zhong, Yao, et al.
Published: (2025)
GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video Generator
by: Zhu, Liyuan, et al.
Published: (2026)
by: Zhu, Liyuan, et al.
Published: (2026)
Lay-A-Scene: Personalized 3D Object Arrangement Using Text-to-Image Priors
by: Rahamim, Ohad, et al.
Published: (2024)
by: Rahamim, Ohad, et al.
Published: (2024)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
by: Chen, Hansheng, et al.
Published: (2024)
by: Chen, Hansheng, et al.
Published: (2024)
Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection
by: Kuang, Zhaonian, et al.
Published: (2026)
by: Kuang, Zhaonian, et al.
Published: (2026)
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
X-Dyna: Expressive Dynamic Human Image Animation
by: Chang, Di, et al.
Published: (2025)
by: Chang, Di, et al.
Published: (2025)
3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation
by: Chen, Hansheng, et al.
Published: (2024)
by: Chen, Hansheng, et al.
Published: (2024)
Garment Particles: A 2D--3D Symmetric Garment Representation for Generation and Editing
by: Nakayama, Kiyohiro, et al.
Published: (2026)
by: Nakayama, Kiyohiro, et al.
Published: (2026)
ReStyle3D: Scene-Level Appearance Transfer with Semantic Correspondences
by: Zhu, Liyuan, et al.
Published: (2025)
by: Zhu, Liyuan, et al.
Published: (2025)
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
by: Chao, Brian, et al.
Published: (2026)
by: Chao, Brian, et al.
Published: (2026)
Orthogonal Adaptation for Modular Customization of Diffusion Models
by: Po, Ryan, et al.
Published: (2023)
by: Po, Ryan, et al.
Published: (2023)
Spectral Progressive Diffusion for Efficient Image and Video Generation
by: Xiao, Howard, et al.
Published: (2026)
by: Xiao, Howard, et al.
Published: (2026)
Policy-based Foveated Imaging and Perception
by: Xiao, Howard, et al.
Published: (2026)
by: Xiao, Howard, et al.
Published: (2026)
MTMMC: A Large-Scale Real-World Multi-Modal Camera Tracking Benchmark
by: Woo, Sanghyun, et al.
Published: (2024)
by: Woo, Sanghyun, et al.
Published: (2024)
ThermalNeRF: Thermal Radiance Fields
by: Lin, Yvette Y., et al.
Published: (2024)
by: Lin, Yvette Y., et al.
Published: (2024)
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
by: Zheng, Yang, et al.
Published: (2025)
by: Zheng, Yang, et al.
Published: (2025)
Locality-Aware Zero-Shot Human-Object Interaction Detection
by: Kim, Sanghyun, et al.
Published: (2025)
by: Kim, Sanghyun, et al.
Published: (2025)
Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting
by: Wang, Weiquan, et al.
Published: (2026)
by: Wang, Weiquan, et al.
Published: (2026)
GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation
by: Xu, Yinghao, et al.
Published: (2024)
by: Xu, Yinghao, et al.
Published: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
Geometric Algebra Planes: Convex Implicit Neural Volumes
by: Sivgin, Irmak, et al.
Published: (2024)
by: Sivgin, Irmak, et al.
Published: (2024)
GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects
by: Shen, Licheng, et al.
Published: (2025)
by: Shen, Licheng, et al.
Published: (2025)
AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
by: Ye, Zi, et al.
Published: (2026)
by: Ye, Zi, et al.
Published: (2026)
Similar Items
-
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
by: Chen, Jiacheng, et al.
Published: (2025) -
GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation
by: Ackermann, Jan, et al.
Published: (2026) -
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024) -
Image Sculpting: Precise Object Editing with 3D Geometry Control
by: Yenphraphai, Jiraphon, et al.
Published: (2024) -
Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark
by: Kuang, Zhengfei, et al.
Published: (2023)