HP-Edit: A Human-Preference Post-Training Framework for Image Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Fan, Wang, Chonghuinan, Lei, Lina, Qiu, Yuping, Xu, Jiaqi, Jiang, Jiaxiu, Qin, Xinran, Chen, Zhikai, Song, Fenglong, Wang, Zhixin, Pei, Renjing, Zuo, Wangmeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918459576680448
author Li, Fan
Wang, Chonghuinan
Lei, Lina
Qiu, Yuping
Xu, Jiaqi
Jiang, Jiaxiu
Qin, Xinran
Chen, Zhikai
Song, Fenglong
Wang, Zhixin
Pei, Renjing
Zuo, Wangmeng
author_facet Li, Fan
Wang, Chonghuinan
Lei, Lina
Qiu, Yuping
Xu, Jiaqi
Jiang, Jiaxiu
Qin, Xinran
Chen, Zhikai
Song, Fenglong
Wang, Zhixin
Pei, Renjing
Zuo, Wangmeng
contents Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Flow-GRPO have further improved generation quality, efficiently applying Reinforcement Learning from Human Feedback (RLHF) to diffusion-based editing remains largely unexplored, due to a lack of scalable human-preference datasets and frameworks tailored to diverse editing needs. To fill this gap, we propose HP-Edit, a post-training framework for Human Preference-aligned Editing, and introduce RealPref-50K, a real-world dataset across eight common tasks and balancing common object editing. Specifically, HP-Edit leverages a small amount of human-preference scoring data and a pretrained visual large language model (VLM) to develop HP-Scorer--an automatic, human preference-aligned evaluator. We then use HP-Scorer both to efficiently build a scalable preference dataset and to serve as the reward function for post-training the editing model. We also introduce RealPref-Bench, a benchmark for evaluating real-world editing performance. Extensive experiments demonstrate that our approach significantly enhances models such as Qwen-Image-Edit-2509, aligning their outputs more closely with human preference.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19406
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HP-Edit: A Human-Preference Post-Training Framework for Image Editing
Li, Fan
Wang, Chonghuinan
Lei, Lina
Qiu, Yuping
Xu, Jiaqi
Jiang, Jiaxiu
Qin, Xinran
Chen, Zhikai
Song, Fenglong
Wang, Zhixin
Pei, Renjing
Zuo, Wangmeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Flow-GRPO have further improved generation quality, efficiently applying Reinforcement Learning from Human Feedback (RLHF) to diffusion-based editing remains largely unexplored, due to a lack of scalable human-preference datasets and frameworks tailored to diverse editing needs. To fill this gap, we propose HP-Edit, a post-training framework for Human Preference-aligned Editing, and introduce RealPref-50K, a real-world dataset across eight common tasks and balancing common object editing. Specifically, HP-Edit leverages a small amount of human-preference scoring data and a pretrained visual large language model (VLM) to develop HP-Scorer--an automatic, human preference-aligned evaluator. We then use HP-Scorer both to efficiently build a scalable preference dataset and to serve as the reward function for post-training the editing model. We also introduce RealPref-Bench, a benchmark for evaluating real-world editing performance. Extensive experiments demonstrate that our approach significantly enhances models such as Qwen-Image-Edit-2509, aligning their outputs more closely with human preference.
title HP-Edit: A Human-Preference Post-Training Framework for Image Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.19406