FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chenxi, Wang, Weijie, Li, Qiang, Lepri, Bruno, Sebe, Nicu, Nie, Weizhi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreeInsert: Personalized Object Insertion with Geometric and Style Control
by: Zhang, Yuhong, et al.
Published: (2025)
by: Zhang, Yuhong, et al.
Published: (2025)
Causal Disentanglement for Robust Long-tail Medical Image Generation
by: Nie, Weizhi, et al.
Published: (2025)
by: Nie, Weizhi, et al.
Published: (2025)
POCI-Diff: Position Objects Consistently and Interactively with 3D-Layout Guided Diffusion
by: Rigo, Andrea, et al.
Published: (2026)
by: Rigo, Andrea, et al.
Published: (2026)
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022)
by: Wang, Weijie, et al.
Published: (2022)
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
by: Rigo, Andrea, et al.
Published: (2025)
by: Rigo, Andrea, et al.
Published: (2025)
SceneExpander: Expanding 3D Scenes with Free-Form Inserted Views
by: He, Zijian, et al.
Published: (2026)
by: He, Zijian, et al.
Published: (2026)
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
Large Language Models for Multimodal Deformable Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
UVMap-ID: A Controllable and Personalized UV Map Generative Model
by: Wang, Weijie, et al.
Published: (2024)
by: Wang, Weijie, et al.
Published: (2024)
Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA
by: Zhu, Xiaorong, et al.
Published: (2026)
by: Zhu, Xiaorong, et al.
Published: (2026)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
Beauty and the Bias: Exploring the Impact of Attractiveness on Multimodal Large Language Models
by: Gulati, Aditya, et al.
Published: (2025)
by: Gulati, Aditya, et al.
Published: (2025)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023)
by: Wang, Weijie, et al.
Published: (2023)
Point2Insert: Video Object Insertion via Sparse Point Guidance
by: Zhou, Yu, et al.
Published: (2026)
by: Zhou, Yu, et al.
Published: (2026)
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
by: Chen, Xinyu, et al.
Published: (2026)
by: Chen, Xinyu, et al.
Published: (2026)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
by: Shahbazi, Mohamad, et al.
Published: (2024)
by: Shahbazi, Mohamad, et al.
Published: (2024)
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
by: Chen, Jinshu, et al.
Published: (2025)
by: Chen, Jinshu, et al.
Published: (2025)
Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
ART3D: 3D Gaussian Splatting for Text-Guided Artistic Scenes Generation
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
by: Ma, Qi, et al.
Published: (2025)
by: Ma, Qi, et al.
Published: (2025)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
EasyInsert: A Data-Efficient and Generalizable Insertion Policy
by: Li, Guanghe, et al.
Published: (2025)
by: Li, Guanghe, et al.
Published: (2025)
Structure Causal Models and LLMs Integration in Medical Visual Question Answering
by: Xu, Zibo, et al.
Published: (2025)
by: Xu, Zibo, et al.
Published: (2025)
Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
by: Zhao, Dong, et al.
Published: (2026)
by: Zhao, Dong, et al.
Published: (2026)
Predicting Covariate-Driven Spatial Deformation for Nonstationary Gaussian Processes
by: Gu, Minghao, et al.
Published: (2026)
by: Gu, Minghao, et al.
Published: (2026)
PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object Detection
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
Organ-Agents: Virtual Human Physiology Simulator via LLMs
by: Chang, Rihao, et al.
Published: (2025)
by: Chang, Rihao, et al.
Published: (2025)
TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment
by: Liu, Jiarun, et al.
Published: (2026)
by: Liu, Jiarun, et al.
Published: (2026)
Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
by: Li, Jinlong, et al.
Published: (2025)
by: Li, Jinlong, et al.
Published: (2025)
DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image
by: Zhao, Qi, et al.
Published: (2025)
by: Zhao, Qi, et al.
Published: (2025)
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
by: Yu, Xiaoxuan, et al.
Published: (2024)
by: Yu, Xiaoxuan, et al.
Published: (2024)
Street Gaussians without 3D Object Tracker
by: Zhang, Ruida, et al.
Published: (2024)
by: Zhang, Ruida, et al.
Published: (2024)
NullFace: Training-Free Localized Face Anonymization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Text‐Guided Interactive Scene Synthesis with Scene Prior Guidance
by: Shaoheng Fang, et al.
Published: (2025)
by: Shaoheng Fang, et al.
Published: (2025)
3D Smoke Scene Reconstruction Guided by Vision Priors from Multimodal Large Language Models
by: Zheng, Xinye, et al.
Published: (2026)
by: Zheng, Xinye, et al.
Published: (2026)
StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians
by: Zhuang, Cailin, et al.
Published: (2025)
by: Zhuang, Cailin, et al.
Published: (2025)
Similar Items
-
FreeInsert: Personalized Object Insertion with Geometric and Style Control
by: Zhang, Yuhong, et al.
Published: (2025) -
Causal Disentanglement for Robust Long-tail Medical Image Generation
by: Nie, Weizhi, et al.
Published: (2025) -
POCI-Diff: Position Objects Consistently and Interactively with 3D-Layout Guided Diffusion
by: Rigo, Andrea, et al.
Published: (2026) -
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022) -
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
by: Wang, Weijie, et al.
Published: (2026)