3D-aware Image Generation and Editing with Multi-modal Conditions
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bo, Li, Yi-ke, He, Zhi-fen, Liu, Bin, Lai, Yun-Kun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification
by: Xie, Ren-Dong, et al.
Published: (2025)
by: Xie, Ren-Dong, et al.
Published: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025)
by: Lu, Jinda, et al.
Published: (2025)
Multi-Reward as Condition for Instruction-based Image Editing
by: Gu, Xin, et al.
Published: (2024)
by: Gu, Xin, et al.
Published: (2024)
FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing
by: Li, Maomao, et al.
Published: (2026)
by: Li, Maomao, et al.
Published: (2026)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
by: Para, Wamiq Reyaz, et al.
Published: (2024)
by: Para, Wamiq Reyaz, et al.
Published: (2024)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
Text-Image Conditioned 3D Generation
by: Cen, Jiazhong, et al.
Published: (2026)
by: Cen, Jiazhong, et al.
Published: (2026)
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
by: He, Yun, et al.
Published: (2025)
by: He, Yun, et al.
Published: (2025)
Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery
by: He, Zhi-Fen, et al.
Published: (2025)
by: He, Zhi-Fen, et al.
Published: (2025)
Dual Diffusion Models for Multi-modal Guided 3D Avatar Generation
by: Li, Hong, et al.
Published: (2026)
by: Li, Hong, et al.
Published: (2026)
R2Human: Real-Time 3D Human Appearance Rendering from a Single Image
by: Yang, Yuanwang, et al.
Published: (2023)
by: Yang, Yuanwang, et al.
Published: (2023)
EIMC: Efficient Instance-aware Multi-modal Collaborative Perception
by: Yang, Kang, et al.
Published: (2026)
by: Yang, Kang, et al.
Published: (2026)
Depth- and Semantics-aware Multi-modal Domain Translation: Generating 3D Panoramic Color Images from LiDAR Point Clouds
by: Cortinhal, Tiago, et al.
Published: (2023)
by: Cortinhal, Tiago, et al.
Published: (2023)
SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model
by: Chen, Guibin, et al.
Published: (2026)
by: Chen, Guibin, et al.
Published: (2026)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing
by: Liu, Xiaolu, et al.
Published: (2026)
by: Liu, Xiaolu, et al.
Published: (2026)
Real-time 3D-aware Portrait Editing from a Single Image
by: Bai, Qingyan, et al.
Published: (2024)
by: Bai, Qingyan, et al.
Published: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
by: Zhang, Kaichen, et al.
Published: (2024)
by: Zhang, Kaichen, et al.
Published: (2024)
OAHuman: Occlusion-Aware 3D Human Reconstruction from Monocular Images
by: Yang, Yuanwang, et al.
Published: (2026)
by: Yang, Yuanwang, et al.
Published: (2026)
Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion
by: Li, Xilai, et al.
Published: (2023)
by: Li, Xilai, et al.
Published: (2023)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
Progressive Multi-modal Conditional Prompt Tuning
by: Qiu, Xiaoyu, et al.
Published: (2024)
by: Qiu, Xiaoyu, et al.
Published: (2024)
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
ShapeUP: Scalable Image-Conditioned 3D Editing
by: Gat, Inbar, et al.
Published: (2026)
by: Gat, Inbar, et al.
Published: (2026)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
MDE-Edit: Masked Dual-Editing for Multi-Object Image Editing via Diffusion Models
by: Zhu, Hongyang, et al.
Published: (2025)
by: Zhu, Hongyang, et al.
Published: (2025)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
by: Li, Zongjian, et al.
Published: (2025)
by: Li, Zongjian, et al.
Published: (2025)
MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing
by: Zhou, Kangneng, et al.
Published: (2023)
by: Zhou, Kangneng, et al.
Published: (2023)
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
by: He, Runze, et al.
Published: (2024)
by: He, Runze, et al.
Published: (2024)
DreamOmni: Unified Image Generation and Editing
by: Xia, Bin, et al.
Published: (2024)
by: Xia, Bin, et al.
Published: (2024)
Generative Video Motion Editing with 3D Point Tracks
by: Lee, Yao-Chih, et al.
Published: (2025)
by: Lee, Yao-Chih, et al.
Published: (2025)
SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation
by: He, Ming, et al.
Published: (2026)
by: He, Ming, et al.
Published: (2026)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
by: Li, Yayuan, et al.
Published: (2024)
by: Li, Yayuan, et al.
Published: (2024)
FineDiffusion: Scaling up Diffusion Models for Fine-grained Image Generation with 10,000 Classes
by: Pan, Ziying, et al.
Published: (2024)
by: Pan, Ziying, et al.
Published: (2024)
Texture-GS: Disentangling the Geometry and Texture for 3D Gaussian Splatting Editing
by: Xu, Tian-Xing, et al.
Published: (2024)
by: Xu, Tian-Xing, et al.
Published: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
by: Yao, Ruilin, et al.
Published: (2024)
by: Yao, Ruilin, et al.
Published: (2024)
HumanCoser: Layered 3D Human Generation via Semantic-Aware Diffusion Model
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
Object-aware Inversion and Reassembly for Image Editing
by: Yang, Zhen, et al.
Published: (2023)
by: Yang, Zhen, et al.
Published: (2023)
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
by: Hu, Zhangchi, et al.
Published: (2026)
by: Hu, Zhangchi, et al.
Published: (2026)
Similar Items
-
Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification
by: Xie, Ren-Dong, et al.
Published: (2025) -
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
by: Lu, Jinda, et al.
Published: (2025) -
Multi-Reward as Condition for Instruction-based Image Editing
by: Gu, Xin, et al.
Published: (2024) -
FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing
by: Li, Maomao, et al.
Published: (2026) -
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
by: Para, Wamiq Reyaz, et al.
Published: (2024)