SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Ziniu, Iscen, Ahmet, Jain, Aashi, Kipf, Thomas, Yue, Yisong, Ross, David A., Schmid, Cordelia, Fathi, Alireza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
von: Caron, Mathilde, et al.
Veröffentlicht: (2024)
SceneCraft: Layout-Guided 3D Scene Generation
von: Yang, Xiuyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiuyu, et al.
Veröffentlicht: (2024)
SceneCrafter: Controllable Multi-View Driving Scene Editing
von: Zhu, Zehao, et al.
Veröffentlicht: (2025)
von: Zhu, Zehao, et al.
Veröffentlicht: (2025)
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
von: Arnab, Anurag, et al.
Veröffentlicht: (2025)
von: Arnab, Anurag, et al.
Veröffentlicht: (2025)
Online 3D Scene Reconstruction Using Neural Object Priors
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
von: Chabal, Thomas, et al.
Veröffentlicht: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
von: Kang, Dahyun, et al.
Veröffentlicht: (2025)
von: Kang, Dahyun, et al.
Veröffentlicht: (2025)
CAViAR: Critic-Augmented Video Agentic Reasoning
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
von: Menon, Sachit, et al.
Veröffentlicht: (2025)
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
von: Huang, Ian, et al.
Veröffentlicht: (2025)
von: Huang, Ian, et al.
Veröffentlicht: (2025)
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
von: Sokar, Ghada, et al.
Veröffentlicht: (2025)
Visual Lexicon: Rich Image Features in Language Space
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
von: Zhang, Dan, et al.
Veröffentlicht: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
von: Uijlings, Jasper, et al.
Veröffentlicht: (2025)
TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
Language-Guided Image Tokenization for Generation
von: Zha, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zha, Kaiwen, et al.
Veröffentlicht: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
von: Garcia, Ricardo, et al.
Veröffentlicht: (2024)
SpaceBlender: Creating Context-Rich Collaborative Spaces Through Generative 3D Scene Blending
von: Numan, Nels, et al.
Veröffentlicht: (2024)
von: Numan, Nels, et al.
Veröffentlicht: (2024)
DORSal: Diffusion for Object-centric Representations of Scenes et al
von: Jabri, Allan, et al.
Veröffentlicht: (2023)
von: Jabri, Allan, et al.
Veröffentlicht: (2023)
SceneGenAgent: Precise Industrial Scene Generation with Coding Agent
von: Xia, Xiao, et al.
Veröffentlicht: (2024)
von: Xia, Xiao, et al.
Veröffentlicht: (2024)
DataSciBench: An LLM Agent Benchmark for Data Science
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
von: Seitzer, Maximilian, et al.
Veröffentlicht: (2023)
von: Seitzer, Maximilian, et al.
Veröffentlicht: (2023)
RoomCraft: Controllable and Complete 3D Indoor Scene Generation
von: Zhou, Mengqi, et al.
Veröffentlicht: (2025)
von: Zhou, Mengqi, et al.
Veröffentlicht: (2025)
Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
von: Cai, Min, et al.
Veröffentlicht: (2024)
von: Cai, Min, et al.
Veröffentlicht: (2024)
Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
von: Light, Jonathan, et al.
Veröffentlicht: (2024)
BrickNet: Graph-Backed Generative Brick Assembly
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
von: Kulits, Peter, et al.
Veröffentlicht: (2026)
Synthesizing Physically Plausible Human Motions in 3D Scenes
von: Pan, Liang, et al.
Veröffentlicht: (2023)
von: Pan, Liang, et al.
Veröffentlicht: (2023)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
von: Wu, Ziyi, et al.
Veröffentlicht: (2024)
von: Wu, Ziyi, et al.
Veröffentlicht: (2024)
Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes
von: Hong, Yuanduo, et al.
Veröffentlicht: (2021)
von: Hong, Yuanduo, et al.
Veröffentlicht: (2021)
Long-Term Human Trajectory Prediction using 3D Dynamic Scene Graphs
von: Gorlo, Nicolas, et al.
Veröffentlicht: (2024)
von: Gorlo, Nicolas, et al.
Veröffentlicht: (2024)
MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
von: Jiang, Hanwen, et al.
Veröffentlicht: (2024)
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
von: von Lützow, Nicolas, et al.
Veröffentlicht: (2026)
von: von Lützow, Nicolas, et al.
Veröffentlicht: (2026)
Beyond Numeric Rewards: In-Context Dueling Bandits with LLM Agents
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
von: Xia, Fanzeng, et al.
Veröffentlicht: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
von: Chen, Shizhe, et al.
Veröffentlicht: (2024)
Self-Evolving Visual Concept Library using Vision-Language Critics
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2025)
SceneScout: Towards AI Agent-driven Access to Street View Imagery for Blind Users
von: Jain, Gaurav, et al.
Veröffentlicht: (2025)
von: Jain, Gaurav, et al.
Veröffentlicht: (2025)
SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation
von: Luo, Jun, et al.
Veröffentlicht: (2026)
von: Luo, Jun, et al.
Veröffentlicht: (2026)
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
von: Lee, Jie-Ying, et al.
Veröffentlicht: (2025)
von: Lee, Jie-Ying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
von: Caron, Mathilde, et al.
Veröffentlicht: (2024) -
Retrieval-Enhanced Contrastive Vision-Text Models
von: Iscen, Ahmet, et al.
Veröffentlicht: (2023) -
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
von: Caron, Mathilde, et al.
Veröffentlicht: (2024) -
SceneCraft: Layout-Guided 3D Scene Generation
von: Yang, Xiuyu, et al.
Veröffentlicht: (2024) -
SceneCrafter: Controllable Multi-View Driving Scene Editing
von: Zhu, Zehao, et al.
Veröffentlicht: (2025)