Saved in:
| Main Authors: | Zhou, Zhimu, Zhao, Yanpeng, Liao, Qiuyu, Zhao, Bo, Ma, Xiaojian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.22868 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025)
by: Wu, Huimin, et al.
Published: (2025)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse
by: Li, Muyao, et al.
Published: (2025)
by: Li, Muyao, et al.
Published: (2025)
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Unsupervised Region-Based Image Editing of Denoising Diffusion Models
by: Li, Zixiang, et al.
Published: (2024)
by: Li, Zixiang, et al.
Published: (2024)
Training-Free Text-Guided Image Editing with Visual Autoregressive Model
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
Contrastive Learning Guided Latent Diffusion Model for Image-to-Image Translation
by: Si, Qi, et al.
Published: (2025)
by: Si, Qi, et al.
Published: (2025)
Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
by: Feng, Zhanbo, et al.
Published: (2023)
by: Feng, Zhanbo, et al.
Published: (2023)
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
by: Yi, Hongzhu, et al.
Published: (2026)
by: Yi, Hongzhu, et al.
Published: (2026)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
by: Zou, Siyu, et al.
Published: (2024)
by: Zou, Siyu, et al.
Published: (2024)
DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing
by: Lyu, Yueming, et al.
Published: (2023)
by: Lyu, Yueming, et al.
Published: (2023)
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
by: Yang, Siwei, et al.
Published: (2025)
by: Yang, Siwei, et al.
Published: (2025)
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
by: Hui, Mude, et al.
Published: (2024)
by: Hui, Mude, et al.
Published: (2024)
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
by: Shi, Zhengpeng, et al.
Published: (2025)
by: Shi, Zhengpeng, et al.
Published: (2025)
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
by: Hou, Jiacheng, et al.
Published: (2026)
by: Hou, Jiacheng, et al.
Published: (2026)
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
by: Yang, Hanqing, et al.
Published: (2026)
by: Yang, Hanqing, et al.
Published: (2026)
ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning
by: Zhao, Yiran, et al.
Published: (2026)
by: Zhao, Yiran, et al.
Published: (2026)
MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation
by: Zhao, Yuan, et al.
Published: (2025)
by: Zhao, Yuan, et al.
Published: (2025)
Probing Human Visual Robustness with Neurally-Guided Deep Neural Networks
by: Shao, Zhenan, et al.
Published: (2024)
by: Shao, Zhenan, et al.
Published: (2024)
Learning Image Priors through Patch-based Diffusion Models for Solving Inverse Problems
by: Hu, Jason, et al.
Published: (2024)
by: Hu, Jason, et al.
Published: (2024)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
by: Lin, Wang, et al.
Published: (2026)
by: Lin, Wang, et al.
Published: (2026)
SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
by: Gu, Jing, et al.
Published: (2024)
by: Gu, Jing, et al.
Published: (2024)
Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing
by: Gao, Bo, et al.
Published: (2024)
by: Gao, Bo, et al.
Published: (2024)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
SINE: SINgle Image Editing with Text-to-Image Diffusion Models
by: Zhang, Zhixing, et al.
Published: (2022)
by: Zhang, Zhixing, et al.
Published: (2022)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Multi-modal Auto-regressive Modeling via Visual Words
by: Peng, Tianshuo, et al.
Published: (2024)
by: Peng, Tianshuo, et al.
Published: (2024)
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
by: Yang, Yujia, et al.
Published: (2026)
by: Yang, Yujia, et al.
Published: (2026)
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
by: Ma, Yiwei, et al.
Published: (2024)
by: Ma, Yiwei, et al.
Published: (2024)
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model
by: Hong, Shibo, et al.
Published: (2026)
by: Hong, Shibo, et al.
Published: (2026)
Visual Prompting for One-shot Controllable Video Editing without Inversion
by: Zhang, Zhengbo, et al.
Published: (2025)
by: Zhang, Zhengbo, et al.
Published: (2025)
Bayesian Optimization for Controlled Image Editing via LLMs
by: Cai, Chengkun, et al.
Published: (2025)
by: Cai, Chengkun, et al.
Published: (2025)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
by: Li, Zejun, et al.
Published: (2025)
by: Li, Zejun, et al.
Published: (2025)
Feedforward 3D Editing via Text-Steerable Image-to-3D
by: Ma, Ziqi, et al.
Published: (2025)
by: Ma, Ziqi, et al.
Published: (2025)
An Interpretable Local Editing Model for Counterfactual Medical Image Generation
by: Min, Hyungi, et al.
Published: (2026)
by: Min, Hyungi, et al.
Published: (2026)
Similar Items
-
NEP: Autoregressive Image Editing via Next Editing Token Prediction
by: Wu, Huimin, et al.
Published: (2025) -
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025) -
JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse
by: Li, Muyao, et al.
Published: (2025) -
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
by: Zhu, Yixuan, et al.
Published: (2025) -
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)