3D-aware Image Generation and Editing with Multi-modal Conditions
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Bo, Li, Yi-ke, He, Zhi-fen, Liu, Bin, Lai, Yun-Kun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification
di: Xie, Ren-Dong, et al.
Pubblicazione: (2025)
di: Xie, Ren-Dong, et al.
Pubblicazione: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
di: Lu, Jinda, et al.
Pubblicazione: (2025)
di: Lu, Jinda, et al.
Pubblicazione: (2025)
Multi-Reward as Condition for Instruction-based Image Editing
di: Gu, Xin, et al.
Pubblicazione: (2024)
di: Gu, Xin, et al.
Pubblicazione: (2024)
FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing
di: Li, Maomao, et al.
Pubblicazione: (2026)
di: Li, Maomao, et al.
Pubblicazione: (2026)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
di: Para, Wamiq Reyaz, et al.
Pubblicazione: (2024)
di: Para, Wamiq Reyaz, et al.
Pubblicazione: (2024)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
di: Wang, Yi, et al.
Pubblicazione: (2023)
di: Wang, Yi, et al.
Pubblicazione: (2023)
Text-Image Conditioned 3D Generation
di: Cen, Jiazhong, et al.
Pubblicazione: (2026)
di: Cen, Jiazhong, et al.
Pubblicazione: (2026)
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
di: He, Yun, et al.
Pubblicazione: (2025)
di: He, Yun, et al.
Pubblicazione: (2025)
Incomplete Multi-Label Image Recognition by Co-learning Semantic-Aware Features and Label Recovery
di: He, Zhi-Fen, et al.
Pubblicazione: (2025)
di: He, Zhi-Fen, et al.
Pubblicazione: (2025)
Dual Diffusion Models for Multi-modal Guided 3D Avatar Generation
di: Li, Hong, et al.
Pubblicazione: (2026)
di: Li, Hong, et al.
Pubblicazione: (2026)
R2Human: Real-Time 3D Human Appearance Rendering from a Single Image
di: Yang, Yuanwang, et al.
Pubblicazione: (2023)
di: Yang, Yuanwang, et al.
Pubblicazione: (2023)
EIMC: Efficient Instance-aware Multi-modal Collaborative Perception
di: Yang, Kang, et al.
Pubblicazione: (2026)
di: Yang, Kang, et al.
Pubblicazione: (2026)
Depth- and Semantics-aware Multi-modal Domain Translation: Generating 3D Panoramic Color Images from LiDAR Point Clouds
di: Cortinhal, Tiago, et al.
Pubblicazione: (2023)
di: Cortinhal, Tiago, et al.
Pubblicazione: (2023)
SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model
di: Chen, Guibin, et al.
Pubblicazione: (2026)
di: Chen, Guibin, et al.
Pubblicazione: (2026)
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition
di: He, Yu, et al.
Pubblicazione: (2026)
di: He, Yu, et al.
Pubblicazione: (2026)
Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing
di: Liu, Xiaolu, et al.
Pubblicazione: (2026)
di: Liu, Xiaolu, et al.
Pubblicazione: (2026)
Real-time 3D-aware Portrait Editing from a Single Image
di: Bai, Qingyan, et al.
Pubblicazione: (2024)
di: Bai, Qingyan, et al.
Pubblicazione: (2024)
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
di: Zhang, Kaichen, et al.
Pubblicazione: (2024)
OAHuman: Occlusion-Aware 3D Human Reconstruction from Monocular Images
di: Yang, Yuanwang, et al.
Pubblicazione: (2026)
di: Yang, Yuanwang, et al.
Pubblicazione: (2026)
Bridging the Gap between Multi-focus and Multi-modal: A Focused Integration Framework for Multi-modal Image Fusion
di: Li, Xilai, et al.
Pubblicazione: (2023)
di: Li, Xilai, et al.
Pubblicazione: (2023)
Disentangling Instruction Influence in Diffusion Transformers for Parallel Multi-Instruction-Guided Image Editing
di: Liu, Hui, et al.
Pubblicazione: (2025)
di: Liu, Hui, et al.
Pubblicazione: (2025)
Progressive Multi-modal Conditional Prompt Tuning
di: Qiu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Qiu, Xiaoyu, et al.
Pubblicazione: (2024)
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
di: Zhang, Yi, et al.
Pubblicazione: (2025)
di: Zhang, Yi, et al.
Pubblicazione: (2025)
ShapeUP: Scalable Image-Conditioned 3D Editing
di: Gat, Inbar, et al.
Pubblicazione: (2026)
di: Gat, Inbar, et al.
Pubblicazione: (2026)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
di: Yuan, Bo, et al.
Pubblicazione: (2024)
di: Yuan, Bo, et al.
Pubblicazione: (2024)
MDE-Edit: Masked Dual-Editing for Multi-Object Image Editing via Diffusion Models
di: Zhu, Hongyang, et al.
Pubblicazione: (2025)
di: Zhu, Hongyang, et al.
Pubblicazione: (2025)
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
di: Li, Zongjian, et al.
Pubblicazione: (2025)
di: Li, Zongjian, et al.
Pubblicazione: (2025)
MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion
di: Liu, Bin, et al.
Pubblicazione: (2026)
di: Liu, Bin, et al.
Pubblicazione: (2026)
MaTe3D: Mask-guided Text-based 3D-aware Portrait Editing
di: Zhou, Kangneng, et al.
Pubblicazione: (2023)
di: Zhou, Kangneng, et al.
Pubblicazione: (2023)
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
di: He, Runze, et al.
Pubblicazione: (2024)
di: He, Runze, et al.
Pubblicazione: (2024)
DreamOmni: Unified Image Generation and Editing
di: Xia, Bin, et al.
Pubblicazione: (2024)
di: Xia, Bin, et al.
Pubblicazione: (2024)
Generative Video Motion Editing with 3D Point Tracks
di: Lee, Yao-Chih, et al.
Pubblicazione: (2025)
di: Lee, Yao-Chih, et al.
Pubblicazione: (2025)
SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation
di: He, Ming, et al.
Pubblicazione: (2026)
di: He, Ming, et al.
Pubblicazione: (2026)
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
di: Li, Yayuan, et al.
Pubblicazione: (2024)
di: Li, Yayuan, et al.
Pubblicazione: (2024)
FineDiffusion: Scaling up Diffusion Models for Fine-grained Image Generation with 10,000 Classes
di: Pan, Ziying, et al.
Pubblicazione: (2024)
di: Pan, Ziying, et al.
Pubblicazione: (2024)
Texture-GS: Disentangling the Geometry and Texture for 3D Gaussian Splatting Editing
di: Xu, Tian-Xing, et al.
Pubblicazione: (2024)
di: Xu, Tian-Xing, et al.
Pubblicazione: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
di: Yao, Ruilin, et al.
Pubblicazione: (2024)
di: Yao, Ruilin, et al.
Pubblicazione: (2024)
HumanCoser: Layered 3D Human Generation via Semantic-Aware Diffusion Model
di: Wang, Yi, et al.
Pubblicazione: (2024)
di: Wang, Yi, et al.
Pubblicazione: (2024)
Object-aware Inversion and Reassembly for Image Editing
di: Yang, Zhen, et al.
Pubblicazione: (2023)
di: Yang, Zhen, et al.
Pubblicazione: (2023)
Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning
di: Hu, Zhangchi, et al.
Pubblicazione: (2026)
di: Hu, Zhangchi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification
di: Xie, Ren-Dong, et al.
Pubblicazione: (2025) -
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
di: Lu, Jinda, et al.
Pubblicazione: (2025) -
Multi-Reward as Condition for Instruction-based Image Editing
di: Gu, Xin, et al.
Pubblicazione: (2024) -
FREE-Edit: Using Editing-aware Injection in Rectified Flow Models for Zero-shot Image-Driven Video Editing
di: Li, Maomao, et al.
Pubblicazione: (2026) -
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
di: Para, Wamiq Reyaz, et al.
Pubblicazione: (2024)