Saved in:
| Main Authors: | Gkotsi, Polytimi Anna, Zadaianchuk, Andrii, Derakhshani, Mohammad Mahdi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.28995 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023)
by: Zadaianchuk, Andrii, et al.
Published: (2023)
Any-Shift Prompting for Generalization over Distributions
by: Xiao, Zehao, et al.
Published: (2024)
by: Xiao, Zehao, et al.
Published: (2024)
NeoBabel: A Multilingual Open Tower for Visual Generation
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025)
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025)
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
by: Zadaianchuk, Andrii, et al.
Published: (2026)
by: Zadaianchuk, Andrii, et al.
Published: (2026)
UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
by: He, Xufan, et al.
Published: (2025)
by: He, Xufan, et al.
Published: (2025)
Temporally Consistent Object-Centric Learning by Contrasting Slots
by: Manasyan, Anna, et al.
Published: (2024)
by: Manasyan, Anna, et al.
Published: (2024)
SALAD: Part-Level Latent Diffusion for 3D Shape Generation and Manipulation
by: Koo, Juil, et al.
Published: (2023)
by: Koo, Juil, et al.
Published: (2023)
Structured 3D Latents for Scalable and Versatile 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2024)
by: Xiang, Jianfeng, et al.
Published: (2024)
PatchAlign3D: Local Feature Alignment for Dense 3D Shape understanding
by: Hadgi, Souhail, et al.
Published: (2026)
by: Hadgi, Souhail, et al.
Published: (2026)
Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination
by: Barcellona, Leonardo, et al.
Published: (2024)
by: Barcellona, Leonardo, et al.
Published: (2024)
A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
by: Xu, Philip
Published: (2025)
by: Xu, Philip
Published: (2025)
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
Morpheus: Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments
by: Zhang, Chenyu, et al.
Published: (2025)
by: Zhang, Chenyu, et al.
Published: (2025)
REPARO: Compositional 3D Assets Generation with Differentiable 3D Layout Alignment
by: Han, Haonan, et al.
Published: (2024)
by: Han, Haonan, et al.
Published: (2024)
SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation
by: Bokhovkin, Alexey, et al.
Published: (2024)
by: Bokhovkin, Alexey, et al.
Published: (2024)
Leveling3D: Leveling Up 3D Reconstruction with Feed-Forward 3D Gaussian Splatting and Geometry-Aware Generation
by: Huang, Yiming, et al.
Published: (2026)
by: Huang, Yiming, et al.
Published: (2026)
LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation
by: Swaminathan, Archana, et al.
Published: (2024)
by: Swaminathan, Archana, et al.
Published: (2024)
Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
by: Gao, Quankai, et al.
Published: (2025)
by: Gao, Quankai, et al.
Published: (2025)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
Rigel3D: Rig-aware Latents for Animation-Ready 3D Asset Generation
by: Chatzis, Nikitas, et al.
Published: (2026)
by: Chatzis, Nikitas, et al.
Published: (2026)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
LPA3D: 3D Room-Level Scene Generation from In-the-Wild Images
by: Yang, Ming-Jia, et al.
Published: (2025)
by: Yang, Ming-Jia, et al.
Published: (2025)
HOTS3D: Hyper-Spherical Optimal Transport for Semantic Alignment of Text-to-3D Generation
by: Li, Zezeng, et al.
Published: (2024)
by: Li, Zezeng, et al.
Published: (2024)
Syn3DTxt: Embedding 3D Cues for Scene Text Generation
by: Hsiung, Li-Syun, et al.
Published: (2025)
by: Hsiung, Li-Syun, et al.
Published: (2025)
From One to More: Contextual Part Latents for 3D Generation
by: Dong, Shaocong, et al.
Published: (2025)
by: Dong, Shaocong, et al.
Published: (2025)
StructLDM: Structured Latent Diffusion for 3D Human Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
by: Yang, Yuanbo, et al.
Published: (2024)
by: Yang, Yuanbo, et al.
Published: (2024)
Native and Compact Structured Latents for 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2025)
by: Xiang, Jianfeng, et al.
Published: (2025)
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces
by: Hadgi, Souhail, et al.
Published: (2025)
by: Hadgi, Souhail, et al.
Published: (2025)
HeadGAP: Few-Shot 3D Head Avatar via Generalizable Gaussian Priors
by: Zheng, Xiaozheng, et al.
Published: (2024)
by: Zheng, Xiaozheng, et al.
Published: (2024)
ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents
by: Chen, Honghua, et al.
Published: (2025)
by: Chen, Honghua, et al.
Published: (2025)
Dual3D: Efficient and Consistent Text-to-3D Generation with Dual-mode Multi-view Latent Diffusion
by: Li, Xinyang, et al.
Published: (2024)
by: Li, Xinyang, et al.
Published: (2024)
JADE: Joint-aware Latent Diffusion for 3D Human Generative Modeling
by: Ji, Haorui, et al.
Published: (2024)
by: Ji, Haorui, et al.
Published: (2024)
Multi-scale Latent Point Consistency Models for 3D Shape Generation
by: Du, Bi'an, et al.
Published: (2024)
by: Du, Bi'an, et al.
Published: (2024)
LN3DIFF++: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation
by: Lan, Yushi, et al.
Published: (2024)
by: Lan, Yushi, et al.
Published: (2024)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
GenAssets: Generating in-the-wild 3D Assets in Latent Space
by: Yang, Ze, et al.
Published: (2026)
by: Yang, Ze, et al.
Published: (2026)
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023)
by: Kabra, Rishabh, et al.
Published: (2023)
Similar Items
-
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023) -
Any-Shift Prompting for Generalization over Distributions
by: Xiao, Zehao, et al.
Published: (2024) -
NeoBabel: A Multilingual Open Tower for Visual Generation
by: Derakhshani, Mohammad Mahdi, et al.
Published: (2025) -
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
by: Zadaianchuk, Andrii, et al.
Published: (2026) -
UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
by: He, Xufan, et al.
Published: (2025)