Learning Continuous 3D Words for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Ta-Ying, Gadelha, Matheus, Groueix, Thibault, Fisher, Matthew, Mech, Radomir, Markham, Andrew, Trigoni, Niki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)
by: Ma, Chenyang, et al.
Published: (2024)
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy
by: Ganeshan, Aditya, et al.
Published: (2024)
by: Ganeshan, Aditya, et al.
Published: (2024)
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026)
by: Saha, Oindrila, et al.
Published: (2026)
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
by: Decatur, Dale, et al.
Published: (2025)
by: Decatur, Dale, et al.
Published: (2025)
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
by: Shin, Sangyun, et al.
Published: (2023)
by: Shin, Sangyun, et al.
Published: (2023)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
by: Saha, Oindrila, et al.
Published: (2025)
by: Saha, Oindrila, et al.
Published: (2025)
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal
by: Aguina-Kang, Rio, et al.
Published: (2026)
by: Aguina-Kang, Rio, et al.
Published: (2026)
Instant3dit: Multiview Inpainting for Fast Editing of 3D Objects
by: Barda, Amir, et al.
Published: (2024)
by: Barda, Amir, et al.
Published: (2024)
Residual Primitive Fitting of 3D Shapes with SuperFrusta
by: Ganeshan, Aditya, et al.
Published: (2025)
by: Ganeshan, Aditya, et al.
Published: (2025)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
Data Factory with Minimal Human Effort Using VLMs
by: Ye, Jiaojiao, et al.
Published: (2025)
by: Ye, Jiaojiao, et al.
Published: (2025)
Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
by: Vankadari, Madhu, et al.
Published: (2024)
by: Vankadari, Madhu, et al.
Published: (2024)
Generative Escher Meshes
by: Aigerman, Noam, et al.
Published: (2023)
by: Aigerman, Noam, et al.
Published: (2023)
VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
by: Zhou, Kaichen, et al.
Published: (2020)
by: Zhou, Kaichen, et al.
Published: (2020)
DynPoint: Dynamic Neural Point For View Synthesis
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
TutteNet: Injective 3D Deformations by Composition of 2D Mesh Deformations
by: Sun, Bo, et al.
Published: (2024)
by: Sun, Bo, et al.
Published: (2024)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
by: Wang, Boyang, et al.
Published: (2025)
by: Wang, Boyang, et al.
Published: (2025)
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
by: Deng, Qianyi, et al.
Published: (2024)
by: Deng, Qianyi, et al.
Published: (2024)
DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction
by: Xu, Qiangeng, et al.
Published: (2019)
by: Xu, Qiangeng, et al.
Published: (2019)
Splat and Replace: 3D Reconstruction with Repetitive Elements
by: Violante, Nicolás, et al.
Published: (2025)
by: Violante, Nicolás, et al.
Published: (2025)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
Pre-training Feature Guided Diffusion Model for Speech Enhancement
by: Yang, Yiyuan, et al.
Published: (2024)
by: Yang, Yiyuan, et al.
Published: (2024)
NOPE: Novel Object Pose Estimation from a Single Image
by: Nguyen, Van Nguyen, et al.
Published: (2023)
by: Nguyen, Van Nguyen, et al.
Published: (2023)
SAMa: Material-aware 3D Selection and Segmentation
by: Fischer, Michael, et al.
Published: (2024)
by: Fischer, Michael, et al.
Published: (2024)
GigaPose: Fast and Robust Novel Object Pose Estimation via One Correspondence
by: Nguyen, Van Nguyen, et al.
Published: (2023)
by: Nguyen, Van Nguyen, et al.
Published: (2023)
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
by: Raji, Fadlullah, et al.
Published: (2026)
by: Raji, Fadlullah, et al.
Published: (2026)
PreciseCam: Precise Camera Control for Text-to-Image Generation
by: Bernal-Berdun, Edurne, et al.
Published: (2025)
by: Bernal-Berdun, Edurne, et al.
Published: (2025)
Material Magic Wand: Material-Aware Grouping of 3D Parts in Untextured Meshes
by: Jain, Umangi, et al.
Published: (2026)
by: Jain, Umangi, et al.
Published: (2026)
MatAtlas: Text-driven Consistent Geometry Texturing and Material Assignment
by: Ceylan, Duygu, et al.
Published: (2024)
by: Ceylan, Duygu, et al.
Published: (2024)
3D-Fixup: Advancing Photo Editing with 3D Priors
by: Cheng, Yen-Chi, et al.
Published: (2025)
by: Cheng, Yen-Chi, et al.
Published: (2025)
Personalized Residuals for Concept-Driven Text-to-Image Generation
by: Ham, Cusuh, et al.
Published: (2024)
by: Ham, Cusuh, et al.
Published: (2024)
Mitigating Cognitive Bias in RLHF by Altering Rationality
by: Horter, Tiffany, et al.
Published: (2026)
by: Horter, Tiffany, et al.
Published: (2026)
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
by: Xu, Shitong, et al.
Published: (2025)
by: Xu, Shitong, et al.
Published: (2025)
PoissonNet: A Local-Global Approach for Learning on Surfaces
by: Maesumi, Arman, et al.
Published: (2025)
by: Maesumi, Arman, et al.
Published: (2025)
WordRobe: Text-Guided Generation of Textured 3D Garments
by: Srivastava, Astitva, et al.
Published: (2024)
by: Srivastava, Astitva, et al.
Published: (2024)
Similar Items
-
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024) -
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024) -
Pattern Analogies: Learning to Perform Programmatic Image Edits by Analogy
by: Ganeshan, Aditya, et al.
Published: (2024) -
3D Space as a Scratchpad for Editable Text-to-Image Generation
by: Saha, Oindrila, et al.
Published: (2026) -
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
by: Decatur, Dale, et al.
Published: (2025)