Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Dian, Zhang, Manyuan, Li, Hongyu, Liu, Hongbo, Zou, Kai, Feng, Kaituo, Li, Hongsheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913153825112064
author Zheng, Dian
Zhang, Manyuan
Li, Hongyu
Liu, Hongbo
Zou, Kai
Feng, Kaituo
Li, Hongsheng
author_facet Zheng, Dian
Zhang, Manyuan
Li, Hongyu
Liu, Hongbo
Zou, Kai
Feng, Kaituo
Li, Hongsheng
contents Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent task conflicts, such strategy requires complex multi-stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade-off rather than true mutual reinforcement. To break this paradigm, we propose Uni-Edit, an intelligent image editing task that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni-Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset. Specifically, we first identify image editing as an inherently ideal general task, as it naturally demands both visual understanding and generation. However, existing editing data relies on simplistic instructions that severely underutilize a model's understanding capacity. To address this, we introduce the first automated and scalable data synthesis pipeline for intelligent editing, transforming diverse VQA data into complex and effective editing instructions with embedded questions and nested logic. This yields Uni-Edit-148k, pairing diverse reasoning-intensive instructions with high-quality edited images. Extensive experiments on BAGEL and Janus-Pro demonstrate that tuning solely on Uni-Edit achieves comprehensive enhancements across all three capabilities without any auxiliary operations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21487
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
Zheng, Dian
Zhang, Manyuan
Li, Hongyu
Liu, Hongbo
Zou, Kai
Feng, Kaituo
Li, Hongsheng
Computer Vision and Pattern Recognition
Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent task conflicts, such strategy requires complex multi-stage pipelines, massive data mixing, and balancing tricks, merely resulting in a performance trade-off rather than true mutual reinforcement. To break this paradigm, we propose Uni-Edit, an intelligent image editing task that serves as the first general task for UMM tuning. Unlike complex mixed pipelines, Uni-Edit improves performance across all three abilities at once using only one task, one training stage, and one dataset. Specifically, we first identify image editing as an inherently ideal general task, as it naturally demands both visual understanding and generation. However, existing editing data relies on simplistic instructions that severely underutilize a model's understanding capacity. To address this, we introduce the first automated and scalable data synthesis pipeline for intelligent editing, transforming diverse VQA data into complex and effective editing instructions with embedded questions and nested logic. This yields Uni-Edit-148k, pairing diverse reasoning-intensive instructions with high-quality edited images. Extensive experiments on BAGEL and Janus-Pro demonstrate that tuning solely on Uni-Edit achieves comprehensive enhancements across all three capabilities without any auxiliary operations.
title Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.21487