Feedforward 3D Editing via Text-Steerable Image-to-3D
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Ziqi, Chen, Hongqiao, Yue, Yisong, Gkioxari, Georgia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Find Any Part in 3D
by: Ma, Ziqi, et al.
Published: (2024)
by: Ma, Ziqi, et al.
Published: (2024)
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025)
by: Sahoo, Aadarsh, et al.
Published: (2025)
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
SAM 3D: 3Dfy Anything in Images
by: SAM 3D Team, et al.
Published: (2025)
by: SAM 3D Team, et al.
Published: (2025)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
by: Kang, Raphi, et al.
Published: (2026)
by: Kang, Raphi, et al.
Published: (2026)
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
by: Meng, Yanxu, et al.
Published: (2025)
by: Meng, Yanxu, et al.
Published: (2025)
LatentEditor: Text Driven Local Editing of 3D Scenes
by: Khalid, Umar, et al.
Published: (2023)
by: Khalid, Umar, et al.
Published: (2023)
Visual Agentic AI for Spatial Reasoning with a Dynamic API
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026)
by: Ma, Ziqi, et al.
Published: (2026)
UniQueR: Unified Query-based Feedforward 3D Reconstruction
by: Peng, Chensheng, et al.
Published: (2026)
by: Peng, Chensheng, et al.
Published: (2026)
Reconstructing Hand-Held Objects in 3D from Images and Videos
by: Wu, Jane, et al.
Published: (2024)
by: Wu, Jane, et al.
Published: (2024)
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning
by: Kumar, Amandeep, et al.
Published: (2024)
by: Kumar, Amandeep, et al.
Published: (2024)
RDSplat: Robust Watermarking for 3D Gaussian Splatting Against 2D and 3D Diffusion Editing
by: Zhao, Longjie, et al.
Published: (2025)
by: Zhao, Longjie, et al.
Published: (2025)
ReplaceAnything3D:Text-Guided 3D Scene Editing with Compositional Neural Radiance Fields
by: Bartrum, Edward, et al.
Published: (2024)
by: Bartrum, Edward, et al.
Published: (2024)
SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
by: Agrawal, Vaibhav, et al.
Published: (2026)
by: Agrawal, Vaibhav, et al.
Published: (2026)
SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors
by: He, Bing, et al.
Published: (2026)
by: He, Bing, et al.
Published: (2026)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
by: Zhang, Frank, et al.
Published: (2024)
by: Zhang, Frank, et al.
Published: (2024)
Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data
by: Xu, Yizhao, et al.
Published: (2026)
by: Xu, Yizhao, et al.
Published: (2026)
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
by: Sahoo, Aadarsh, et al.
Published: (2026)
by: Sahoo, Aadarsh, et al.
Published: (2026)
3D MRI Image Pretraining via Controllable 2D Slice Navigation Task
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
SwinTF3D: A Lightweight Multimodal Fusion Approach for Text-Guided 3D Medical Image Segmentation
by: Khan, Hasan Faraz, et al.
Published: (2025)
by: Khan, Hasan Faraz, et al.
Published: (2025)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
Steerable Visual Representations
by: Ruthardt, Jona, et al.
Published: (2026)
by: Ruthardt, Jona, et al.
Published: (2026)
DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
Proc3D: Procedural 3D Generation and Parametric Editing of 3D Shapes with Large Language Models
by: Raji, Fadlullah, et al.
Published: (2026)
by: Raji, Fadlullah, et al.
Published: (2026)
Text-Driven Image Editing via Learnable Regions
by: Lin, Yuanze, et al.
Published: (2023)
by: Lin, Yuanze, et al.
Published: (2023)
Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
by: Kim, Hayeon, et al.
Published: (2025)
by: Kim, Hayeon, et al.
Published: (2025)
Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
by: Dunlop, Connor, et al.
Published: (2025)
by: Dunlop, Connor, et al.
Published: (2025)
Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
ICE-G: Image Conditional Editing of 3D Gaussian Splats
by: Jaganathan, Vishnu, et al.
Published: (2024)
by: Jaganathan, Vishnu, et al.
Published: (2024)
Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion
by: Tang, Zhenggang, et al.
Published: (2026)
by: Tang, Zhenggang, et al.
Published: (2026)
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
DragText: Rethinking Text Embedding in Point-based Image Editing
by: Choi, Gayoon, et al.
Published: (2024)
by: Choi, Gayoon, et al.
Published: (2024)
Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
by: Ye, Chongjie, et al.
Published: (2026)
by: Ye, Chongjie, et al.
Published: (2026)
Similar Items
-
Find Any Part in 3D
by: Ma, Ziqi, et al.
Published: (2024) -
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025) -
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
by: Marsili, Damiano, et al.
Published: (2025) -
SAM 3D: 3Dfy Anything in Images
by: SAM 3D Team, et al.
Published: (2025) -
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)