Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
Fuente:
arXiv
Saved in:
| Main Authors: | Ci, En, Guan, Shanyan, Ge, Yanhao, Zhang, Yilin, Li, Wei, Zhang, Zhenyu, Yang, Jian, Tai, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
by: Chen, Zhizhou, et al.
Published: (2026)
by: Chen, Zhizhou, et al.
Published: (2026)
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
by: Zhao, Chen, et al.
Published: (2025)
by: Zhao, Chen, et al.
Published: (2025)
RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations
by: Ge, Yanhao, et al.
Published: (2026)
by: Ge, Yanhao, et al.
Published: (2026)
HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation
by: Guan, Shanyan, et al.
Published: (2024)
by: Guan, Shanyan, et al.
Published: (2024)
PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
by: Tian, Feng, et al.
Published: (2024)
by: Tian, Feng, et al.
Published: (2024)
ACE-LoRA: Adaptive Orthogonal Decoupling for Continual Image Editing
by: Liu, Yuehao, et al.
Published: (2026)
by: Liu, Yuehao, et al.
Published: (2026)
Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models
by: Liu, Yuehao, et al.
Published: (2026)
by: Liu, Yuehao, et al.
Published: (2026)
NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
by: Zhao, Yanpeng, et al.
Published: (2025)
by: Zhao, Yanpeng, et al.
Published: (2025)
Neural Material Adaptor for Visual Grounding of Intrinsic Dynamics
by: Cao, Junyi, et al.
Published: (2024)
by: Cao, Junyi, et al.
Published: (2024)
Guiding a Diffusion Model by Swapping Its Tokens
by: Zhang, Weijia, et al.
Published: (2026)
by: Zhang, Weijia, et al.
Published: (2026)
Don't Forget your Inverse DDIM for Image Editing
by: Gomez-Trenado, Guillermo, et al.
Published: (2025)
by: Gomez-Trenado, Guillermo, et al.
Published: (2025)
DiffFAE: Advancing High-fidelity One-shot Facial Appearance Editing with Space-sensitive Customization and Semantic Preservation
by: Wang, Qilin, et al.
Published: (2024)
by: Wang, Qilin, et al.
Published: (2024)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
Describing Differences in Image Sets with Natural Language
by: Dunlap, Lisa, et al.
Published: (2023)
by: Dunlap, Lisa, et al.
Published: (2023)
Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification
by: Yang, Yuting, et al.
Published: (2026)
by: Yang, Yuting, et al.
Published: (2026)
StrandHead: Text to Hair-Disentangled 3D Head Avatars Using Human-Centric Priors
by: Sun, Xiaokun, et al.
Published: (2024)
by: Sun, Xiaokun, et al.
Published: (2024)
DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies
by: Wang, Renke, et al.
Published: (2026)
by: Wang, Renke, et al.
Published: (2026)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
Layton: Latent Consistency Tokenizer for 1024-pixel Image Reconstruction and Generation by 256 Tokens
by: Xie, Qingsong, et al.
Published: (2025)
by: Xie, Qingsong, et al.
Published: (2025)
RadDiff: Describing Differences in Radiology Image Sets with Natural Language
by: Shen, Xiaoxian, et al.
Published: (2026)
by: Shen, Xiaoxian, et al.
Published: (2026)
Don't let the information slip away
by: Li, Taozhe, et al.
Published: (2026)
by: Li, Taozhe, et al.
Published: (2026)
Don't Look into the Dark: Latent Codes for Pluralistic Image Inpainting
by: Chen, Haiwei, et al.
Published: (2024)
by: Chen, Haiwei, et al.
Published: (2024)
When Text and Images Don't Mix: Bias-Correcting Language-Image Similarity Scores for Anomaly Detection
by: Goodge, Adam, et al.
Published: (2024)
by: Goodge, Adam, et al.
Published: (2024)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
by: Sun, Xiaokun, et al.
Published: (2026)
by: Sun, Xiaokun, et al.
Published: (2026)
DreamBarbie: Text to Barbie-Style 3D Avatars
by: Sun, Xiaokun, et al.
Published: (2024)
by: Sun, Xiaokun, et al.
Published: (2024)
ScribbleSense: Generative Scribble-Based Texture Editing with Intent Prediction
by: Zhang, Yudi, et al.
Published: (2026)
by: Zhang, Yudi, et al.
Published: (2026)
Don't Fear Peculiar Activation Functions: EUAF and Beyond
by: Wang, Qianchao, et al.
Published: (2024)
by: Wang, Qianchao, et al.
Published: (2024)
Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory Guidance
by: Yang, Haijie, et al.
Published: (2025)
by: Yang, Haijie, et al.
Published: (2025)
RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion
by: Yan, Zhiqiang, et al.
Published: (2023)
by: Yan, Zhiqiang, et al.
Published: (2023)
Semantic Granularity Navigation in Image Editing
by: Lu, Liangsi, et al.
Published: (2026)
by: Lu, Liangsi, et al.
Published: (2026)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
Large Pre-Training Datasets Don't Always Guarantee Robustness after Fine-Tuning
by: Hwang, Jaedong, et al.
Published: (2024)
by: Hwang, Jaedong, et al.
Published: (2024)
FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
3D Gaussian Editing with A Single Image
by: Luo, Guan, et al.
Published: (2024)
by: Luo, Guan, et al.
Published: (2024)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Describe Anything in Medical Images
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
DescribeEarth: Describe Anything for Remote Sensing Images
by: Li, Kaiyu, et al.
Published: (2025)
by: Li, Kaiyu, et al.
Published: (2025)
Similar Items
-
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
by: Chen, Zhizhou, et al.
Published: (2026) -
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
by: Zhao, Chen, et al.
Published: (2025) -
RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations
by: Ge, Yanhao, et al.
Published: (2026) -
HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation
by: Guan, Shanyan, et al.
Published: (2024) -
PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
by: Tian, Feng, et al.
Published: (2024)