VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Meng, Yang, Senqiao, Che, Haoxuan, Zhang, Suiyun, Zhang, Xichen, Yu, Shaozuo, Gui, Haokun, Rao, Zhefan, Tu, Dandan, Liu, Rui, Jia, Jiaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910011946434560
author Chu, Meng
Yang, Senqiao
Che, Haoxuan
Zhang, Suiyun
Zhang, Xichen
Yu, Shaozuo
Gui, Haokun
Rao, Zhefan
Tu, Dandan
Liu, Rui
Jia, Jiaya
author_facet Chu, Meng
Yang, Senqiao
Che, Haoxuan
Zhang, Suiyun
Zhang, Xichen
Yu, Shaozuo
Gui, Haokun
Rao, Zhefan
Tu, Dandan
Liu, Rui
Jia, Jiaya
contents Generative models can now produce photorealistic imagery, yet they still struggle with the long, multi-goal prompts that professional designers issue. To expose this gap and better evaluate models' performance in real-world settings, we introduce Long Goal Bench (LGBench), a 2,000-task suite (1,000 T2I and 1,000 I2I) whose average instruction contains 18 to 22 tightly coupled goals spanning global layout, local object placement, typography, and logo fidelity. We find that even state-of-the-art models satisfy fewer than 72 percent of the goals and routinely miss localized edits, confirming the brittleness of current pipelines. To address this, we present VisionDirector, a training-free vision-language supervisor that (i) extracts structured goals from long instructions, (ii) dynamically decides between one-shot generation and staged edits, (iii) runs micro-grid sampling with semantic verification and rollback after every edit, and (iv) logs goal-level rewards. We further fine-tune the planner with Group Relative Policy Optimization, yielding shorter edit trajectories (3.1 versus 4.2 steps) and stronger alignment. VisionDirector achieves new state of the art on GenEval (plus 7 percent overall) and ImgEdit (plus 0.07 absolute) while producing consistent qualitative improvements on typography, multi-object scenes, and pose editing.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19243
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
Chu, Meng
Yang, Senqiao
Che, Haoxuan
Zhang, Suiyun
Zhang, Xichen
Yu, Shaozuo
Gui, Haokun
Rao, Zhefan
Tu, Dandan
Liu, Rui
Jia, Jiaya
Computer Vision and Pattern Recognition
Generative models can now produce photorealistic imagery, yet they still struggle with the long, multi-goal prompts that professional designers issue. To expose this gap and better evaluate models' performance in real-world settings, we introduce Long Goal Bench (LGBench), a 2,000-task suite (1,000 T2I and 1,000 I2I) whose average instruction contains 18 to 22 tightly coupled goals spanning global layout, local object placement, typography, and logo fidelity. We find that even state-of-the-art models satisfy fewer than 72 percent of the goals and routinely miss localized edits, confirming the brittleness of current pipelines. To address this, we present VisionDirector, a training-free vision-language supervisor that (i) extracts structured goals from long instructions, (ii) dynamically decides between one-shot generation and staged edits, (iii) runs micro-grid sampling with semantic verification and rollback after every edit, and (iv) logs goal-level rewards. We further fine-tune the planner with Group Relative Policy Optimization, yielding shorter edit trajectories (3.1 versus 4.2 steps) and stronger alignment. VisionDirector achieves new state of the art on GenEval (plus 7 percent overall) and ImgEdit (plus 0.07 absolute) while producing consistent qualitative improvements on typography, multi-object scenes, and pose editing.
title VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19243