Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Ziyi, Rubanova, Yulia, Kabra, Rishabh, Hudson, Drew A., Gilitschenski, Igor, Aytar, Yusuf, van Steenkiste, Sjoerd, Allen, Kelsey R., Kipf, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Neural USD: An object-centric framework for iterative editing and control
di: Escontrela, Alejandro, et al.
Pubblicazione: (2025)
di: Escontrela, Alejandro, et al.
Pubblicazione: (2025)
How to Spin an Object: First, Get the Shape Right
di: Kabra, Rishabh, et al.
Pubblicazione: (2024)
di: Kabra, Rishabh, et al.
Pubblicazione: (2024)
DORSal: Diffusion for Object-centric Representations of Scenes et al
di: Jabri, Allan, et al.
Pubblicazione: (2023)
di: Jabri, Allan, et al.
Pubblicazione: (2023)
Moving Off-the-Grid: Scene-Grounded Video Representations
di: van Steenkiste, Sjoerd, et al.
Pubblicazione: (2024)
di: van Steenkiste, Sjoerd, et al.
Pubblicazione: (2024)
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
di: Seitzer, Maximilian, et al.
Pubblicazione: (2023)
di: Seitzer, Maximilian, et al.
Pubblicazione: (2023)
Direct Motion Models for Assessing Generated Videos
di: Allen, Kelsey, et al.
Pubblicazione: (2025)
di: Allen, Kelsey, et al.
Pubblicazione: (2025)
Scaling Face Interaction Graph Networks to Real World Scenes
di: Lopez-Guevara, Tatiana, et al.
Pubblicazione: (2024)
di: Lopez-Guevara, Tatiana, et al.
Pubblicazione: (2024)
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
di: Qiu, Linlu, et al.
Pubblicazione: (2025)
di: Qiu, Linlu, et al.
Pubblicazione: (2025)
How Does Code Pretraining Affect Language Model Task Performance?
di: Petty, Jackson, et al.
Pubblicazione: (2024)
di: Petty, Jackson, et al.
Pubblicazione: (2024)
CTRL-D: Controllable Dynamic 3D Scene Editing with Personalized 2D Diffusion
di: He, Kai, et al.
Pubblicazione: (2024)
di: He, Kai, et al.
Pubblicazione: (2024)
Learning rigid-body simulators over implicit shapes for large-scale scenes and vision
di: Rubanova, Yulia, et al.
Pubblicazione: (2024)
di: Rubanova, Yulia, et al.
Pubblicazione: (2024)
LEOD: Label-Efficient Object Detection for Event Cameras
di: Wu, Ziyi, et al.
Pubblicazione: (2023)
di: Wu, Ziyi, et al.
Pubblicazione: (2023)
TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
di: Mohammadi, Mohammad, et al.
Pubblicazione: (2025)
di: Mohammadi, Mohammad, et al.
Pubblicazione: (2025)
Leveraging VLM-Based Pipelines to Annotate 3D Objects
di: Kabra, Rishabh, et al.
Pubblicazione: (2023)
di: Kabra, Rishabh, et al.
Pubblicazione: (2023)
SPAD : Spatially Aware Multiview Diffusers
di: Kant, Yash, et al.
Pubblicazione: (2024)
di: Kant, Yash, et al.
Pubblicazione: (2024)
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
di: Kabra, Rishabh, et al.
Pubblicazione: (2026)
di: Kabra, Rishabh, et al.
Pubblicazione: (2026)
FlexCap: Describe Anything in Images in Controllable Detail
di: Dwibedi, Debidatta, et al.
Pubblicazione: (2024)
di: Dwibedi, Debidatta, et al.
Pubblicazione: (2024)
A Systematic Comparison of Syllogistic Reasoning in Humans and Language Models
di: Eisape, Tiwalayo, et al.
Pubblicazione: (2023)
di: Eisape, Tiwalayo, et al.
Pubblicazione: (2023)
The Impact of Depth on Compositional Generalization in Transformer Language Models
di: Petty, Jackson, et al.
Pubblicazione: (2023)
di: Petty, Jackson, et al.
Pubblicazione: (2023)
Corra: Correlation-Aware Column Compression
di: Liu, Hanwen, et al.
Pubblicazione: (2024)
di: Liu, Hanwen, et al.
Pubblicazione: (2024)
SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
di: Namekata, Koichi, et al.
Pubblicazione: (2024)
di: Namekata, Koichi, et al.
Pubblicazione: (2024)
Towards Unsupervised Blind Face Restoration using Diffusion Prior
di: Kuai, Tianshu, et al.
Pubblicazione: (2024)
di: Kuai, Tianshu, et al.
Pubblicazione: (2024)
From Image to Video: An Empirical Study of Diffusion Representations
di: Vélez, Pedro, et al.
Pubblicazione: (2025)
di: Vélez, Pedro, et al.
Pubblicazione: (2025)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
di: Dwibedi, Debidatta, et al.
Pubblicazione: (2024)
di: Dwibedi, Debidatta, et al.
Pubblicazione: (2024)
Watch Your Steps: Local Image and Scene Editing by Text Instructions
di: Mirzaei, Ashkan, et al.
Pubblicazione: (2023)
di: Mirzaei, Ashkan, et al.
Pubblicazione: (2023)
RefFusion: Reference Adapted Diffusion Models for 3D Scene Inpainting
di: Mirzaei, Ashkan, et al.
Pubblicazione: (2024)
di: Mirzaei, Ashkan, et al.
Pubblicazione: (2024)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
di: Wu, Ziyi, et al.
Pubblicazione: (2025)
di: Wu, Ziyi, et al.
Pubblicazione: (2025)
Image Synthesis with Class-Aware Semantic Diffusion Models for Surgical Scene Segmentation
di: Zhou, Yihang, et al.
Pubblicazione: (2024)
di: Zhou, Yihang, et al.
Pubblicazione: (2024)
Scaling 4D Representations
di: Carreira, João, et al.
Pubblicazione: (2024)
di: Carreira, João, et al.
Pubblicazione: (2024)
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
di: Nayak, Shravan, et al.
Pubblicazione: (2025)
di: Nayak, Shravan, et al.
Pubblicazione: (2025)
S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control
di: Liu, Xudong, et al.
Pubblicazione: (2025)
di: Liu, Xudong, et al.
Pubblicazione: (2025)
Material Magic Wand: Material-Aware Grouping of 3D Parts in Untextured Meshes
di: Jain, Umangi, et al.
Pubblicazione: (2026)
di: Jain, Umangi, et al.
Pubblicazione: (2026)
Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
di: Zheng, Shuhong, et al.
Pubblicazione: (2025)
di: Zheng, Shuhong, et al.
Pubblicazione: (2025)
GaussianCut: Interactive segmentation via graph cut for 3D Gaussian Splatting
di: Jain, Umangi, et al.
Pubblicazione: (2024)
di: Jain, Umangi, et al.
Pubblicazione: (2024)
GeoMatch++: Morphology Conditioned Geometry Matching for Multi-Embodiment Grasping
di: Wei, Yunze, et al.
Pubblicazione: (2024)
di: Wei, Yunze, et al.
Pubblicazione: (2024)
EventSplat: 3D Gaussian Splatting from Moving Event Cameras for Real-time Rendering
di: Yura, Toshiya, et al.
Pubblicazione: (2024)
di: Yura, Toshiya, et al.
Pubblicazione: (2024)
OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts
di: Xiao, Shiting, et al.
Pubblicazione: (2025)
di: Xiao, Shiting, et al.
Pubblicazione: (2025)
MetaFind: Scene-Aware 3D Asset Retrieval for Coherent Metaverse Scene Generation
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
di: Dwibedi, Debidatta, et al.
Pubblicazione: (2024)
di: Dwibedi, Debidatta, et al.
Pubblicazione: (2024)
Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
di: Biswas, Sandika, et al.
Pubblicazione: (2025)
di: Biswas, Sandika, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Neural USD: An object-centric framework for iterative editing and control
di: Escontrela, Alejandro, et al.
Pubblicazione: (2025) -
How to Spin an Object: First, Get the Shape Right
di: Kabra, Rishabh, et al.
Pubblicazione: (2024) -
DORSal: Diffusion for Object-centric Representations of Scenes et al
di: Jabri, Allan, et al.
Pubblicazione: (2023) -
Moving Off-the-Grid: Scene-Grounded Video Representations
di: van Steenkiste, Sjoerd, et al.
Pubblicazione: (2024) -
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos
di: Seitzer, Maximilian, et al.
Pubblicazione: (2023)