An Item is Worth a Prompt: Versatile Image Editing with Disentangled Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Aosong, Qiu, Weikang, Bai, Jinbin, Zhang, Xiao, Dong, Zhen, Zhou, Kaicheng, Ying, Rex, Tassiulas, Leandros |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Integrating View Conditions for Image Synthesis
von: Bai, Jinbin, et al.
Veröffentlicht: (2023)
von: Bai, Jinbin, et al.
Veröffentlicht: (2023)
Long Sequence Modeling with Attention Tensorization: From Sequence to Tensor Learning
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
Personalized Safety Alignment for Text-to-Image Diffusion Models
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
von: Qiu, Weikang, et al.
Veröffentlicht: (2026)
A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
von: Zhuang, Junhao, et al.
Veröffentlicht: (2023)
von: Zhuang, Junhao, et al.
Veröffentlicht: (2023)
MindLLM: A Subject-Agnostic and Versatile Model for fMRI-to-Text Decoding
von: Qiu, Weikang, et al.
Veröffentlicht: (2025)
von: Qiu, Weikang, et al.
Veröffentlicht: (2025)
DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
von: He, Jing, et al.
Veröffentlicht: (2024)
von: He, Jing, et al.
Veröffentlicht: (2024)
Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs
von: Liu, Jinming, et al.
Veröffentlicht: (2024)
von: Liu, Jinming, et al.
Veröffentlicht: (2024)
Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
von: Feng, Zhanbo, et al.
Veröffentlicht: (2023)
von: Feng, Zhanbo, et al.
Veröffentlicht: (2023)
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
von: Yang, Yitong, et al.
Veröffentlicht: (2024)
von: Yang, Yitong, et al.
Veröffentlicht: (2024)
TextureDiffusion: Target Prompt Disentangled Editing for Various Texture Transfer
von: Su, Zihan, et al.
Veröffentlicht: (2024)
von: Su, Zihan, et al.
Veröffentlicht: (2024)
DesignEdit: Multi-Layered Latent Decomposition and Fusion for Unified & Accurate Image Editing
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
von: Jia, Yueru, et al.
Veröffentlicht: (2024)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
von: Bai, Jinbin, et al.
Veröffentlicht: (2024)
Image-to-Image Translation with Disentangled Latent Vectors for Face Editing
von: Dalva, Yusuf, et al.
Veröffentlicht: (2023)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2023)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
von: Luo, Haochen, et al.
Veröffentlicht: (2024)
von: Luo, Haochen, et al.
Veröffentlicht: (2024)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
DisControlFace: Adding Disentangled Control to Diffusion Autoencoder for One-shot Explicit Facial Image Editing
von: Jia, Haozhe, et al.
Veröffentlicht: (2023)
von: Jia, Haozhe, et al.
Veröffentlicht: (2023)
EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
von: Litman, Yehonathan, et al.
Veröffentlicht: (2026)
von: Litman, Yehonathan, et al.
Veröffentlicht: (2026)
Training-free Geometric Image Editing on Diffusion Models
von: Zhu, Hanshen, et al.
Veröffentlicht: (2025)
von: Zhu, Hanshen, et al.
Veröffentlicht: (2025)
TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval
von: Chen, Jialin, et al.
Veröffentlicht: (2025)
von: Chen, Jialin, et al.
Veröffentlicht: (2025)
Efficient High-Resolution Time Series Classification via Attention Kronecker Decomposition
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
von: Feng, Aosong, et al.
Veröffentlicht: (2024)
Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing
von: Xu, Pengcheng, et al.
Veröffentlicht: (2024)
von: Xu, Pengcheng, et al.
Veröffentlicht: (2024)
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing
von: Wei, Tianyi, et al.
Veröffentlicht: (2025)
von: Wei, Tianyi, et al.
Veröffentlicht: (2025)
MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
von: Zhou, Donghao, et al.
Veröffentlicht: (2024)
von: Zhou, Donghao, et al.
Veröffentlicht: (2024)
Co-Seg++: Mutual Prompt-Guided Collaborative Learning for Versatile Medical Segmentation
von: Xu, Qing, et al.
Veröffentlicht: (2025)
von: Xu, Qing, et al.
Veröffentlicht: (2025)
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
von: Li, Maomao, et al.
Veröffentlicht: (2023)
von: Li, Maomao, et al.
Veröffentlicht: (2023)
Images are Worth Variable Length of Representations
von: Mao, Lingjun, et al.
Veröffentlicht: (2025)
von: Mao, Lingjun, et al.
Veröffentlicht: (2025)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting
von: Zhang, Qihang, et al.
Veröffentlicht: (2024)
von: Zhang, Qihang, et al.
Veröffentlicht: (2024)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
von: Wu, Chen, et al.
Veröffentlicht: (2024)
von: Wu, Chen, et al.
Veröffentlicht: (2024)
Dynamic Motion Blending for Versatile Motion Editing
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
von: Jiang, Nan, et al.
Veröffentlicht: (2025)
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance
von: Lin, Yiqi, et al.
Veröffentlicht: (2026)
von: Lin, Yiqi, et al.
Veröffentlicht: (2026)
TalkPhoto: A Versatile Training-Free Conversational Assistant for Intelligent Image Editing
von: Hu, Yujie, et al.
Veröffentlicht: (2026)
von: Hu, Yujie, et al.
Veröffentlicht: (2026)
Versatile Transition Generation with Image-to-Video Diffusion
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
APE: Agentic Prompt Enhancer for Image Generation and Editing
von: Huang, Zijian, et al.
Veröffentlicht: (2026)
von: Huang, Zijian, et al.
Veröffentlicht: (2026)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
Symmetry Understanding of 3D Shapes via Chirality Disentanglement
von: Wang, Weikang, et al.
Veröffentlicht: (2025)
von: Wang, Weikang, et al.
Veröffentlicht: (2025)
Masked Generative Transformer Is What You Need for Image Editing
von: Chow, Wei, et al.
Veröffentlicht: (2026)
von: Chow, Wei, et al.
Veröffentlicht: (2026)
AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
von: Niu, Yunfang, et al.
Veröffentlicht: (2024)
von: Niu, Yunfang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Integrating View Conditions for Image Synthesis
von: Bai, Jinbin, et al.
Veröffentlicht: (2023) -
Long Sequence Modeling with Attention Tensorization: From Sequence to Tensor Learning
von: Feng, Aosong, et al.
Veröffentlicht: (2024) -
Personalized Safety Alignment for Text-to-Image Diffusion Models
von: Lei, Yu, et al.
Veröffentlicht: (2025) -
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
von: Qiu, Weikang, et al.
Veröffentlicht: (2026) -
A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
von: Zhuang, Junhao, et al.
Veröffentlicht: (2023)