SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ren, Tianfei, Yan, Zhipeng, Zhao, Yiming, Fang, Zhen, Zeng, Yu, Zhang, Guohui, Xu, Hang, Ma, Xiaoxiao, Huang, Shiting, Xu, Ke, Huang, Wenxuan, Wang, Lionel Z., Chen, Lin, Chen, Zehui, Huang, Jie, Zhao, Feng
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915994256015360
author Ren, Tianfei
Yan, Zhipeng
Zhao, Yiming
Fang, Zhen
Zeng, Yu
Zhang, Guohui
Xu, Hang
Ma, Xiaoxiao
Huang, Shiting
Xu, Ke
Huang, Wenxuan
Wang, Lionel Z.
Chen, Lin
Chen, Zehui
Huang, Jie
Zhao, Feng
author_facet Ren, Tianfei
Yan, Zhipeng
Zhao, Yiming
Fang, Zhen
Zeng, Yu
Zhang, Guohui
Xu, Hang
Ma, Xiaoxiao
Huang, Shiting
Xu, Ke
Huang, Wenxuan
Wang, Lionel Z.
Chen, Lin
Chen, Zehui
Huang, Jie
Zhao, Feng
contents While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many requirements must be tracked across grounding, generation, and verification. We refer to these requirements as semantic commitments and formalize their lifecycle discontinuity as the Conceptual Rift, where commitments may be locally resolved or checked but fail to remain identifiable as the same operational units throughout the generation lifecycle. To address this, we propose SCOPE, a specification-guided skill orchestration framework that maintains semantic commitments in an evolving structured specification and conditionally invokes retrieval, reasoning, and repair skills around unresolved or violated commitments. To evaluate commitment-level intent realization, we introduce Gen-Arena, a human-annotated benchmark with entity- and constraint-level specifications, together with Entity-Gated Intent Pass Rate (EGIP), a strict entity-first pass criterion. SCOPE substantially outperforms all evaluated baselines on Gen-Arena, achieving 0.60 EGIP, and further achieves strong results on WISE-V (0.907) and MindBench (0.61), demonstrating the effectiveness of persistent commitment tracking for complex image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08043
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
Ren, Tianfei
Yan, Zhipeng
Zhao, Yiming
Fang, Zhen
Zeng, Yu
Zhang, Guohui
Xu, Hang
Ma, Xiaoxiao
Huang, Shiting
Xu, Ke
Huang, Wenxuan
Wang, Lionel Z.
Chen, Lin
Chen, Zehui
Huang, Jie
Zhao, Feng
Computer Vision and Pattern Recognition
Artificial Intelligence
While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many requirements must be tracked across grounding, generation, and verification. We refer to these requirements as semantic commitments and formalize their lifecycle discontinuity as the Conceptual Rift, where commitments may be locally resolved or checked but fail to remain identifiable as the same operational units throughout the generation lifecycle. To address this, we propose SCOPE, a specification-guided skill orchestration framework that maintains semantic commitments in an evolving structured specification and conditionally invokes retrieval, reasoning, and repair skills around unresolved or violated commitments. To evaluate commitment-level intent realization, we introduce Gen-Arena, a human-annotated benchmark with entity- and constraint-level specifications, together with Entity-Gated Intent Pass Rate (EGIP), a strict entity-first pass criterion. SCOPE substantially outperforms all evaluated baselines on Gen-Arena, achieving 0.60 EGIP, and further achieves strong results on WISE-V (0.907) and MindBench (0.61), demonstrating the effectiveness of persistent commitment tracking for complex image generation.
title SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.08043