AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Hanjun, Huang, Zhimu, Chung, Sylvia, Wang, Yiran, Jin, Yingbin, Li, Jialin, Li, Jiang, Li, Xinfeng, Salam, Hanan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910246497157120
author Luo, Hanjun
Huang, Zhimu
Chung, Sylvia
Wang, Yiran
Jin, Yingbin
Li, Jialin
Li, Jiang
Li, Xinfeng
Salam, Hanan
author_facet Luo, Hanjun
Huang, Zhimu
Chung, Sylvia
Wang, Yiran
Jin, Yingbin
Li, Jialin
Li, Jiang
Li, Xinfeng
Salam, Hanan
contents Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I models, leaving the prompting proficiency of this upstream component entirely unmeasured. We introduce AtelierEval, the first unified benchmark that quantifies prompting proficiency across 360 expert-crafted tasks. Grounded in a cognitive view, it spans three task categories and instantiates tasks using a taxonomy of real-world challenges, with a dual interface for both humans and MLLMs. To enable scalable and reliable evaluation, we propose AtelierJudge, a skill-based, memory-augmented agentic evaluator. It produces subjective and objective scores for prompt-image pairs, achieving a Spearman correlation of 0.79 with human experts, approaching human performance. Extensive experiments benchmark 8 MLLMs against 48 human users across 4 T2I backends, validate AtelierEval as a robust diagnostic tool, and reveal the superiority of mimicry over planning, advocating for an image-augmented direction for future prompters. Our work is released to support future research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22645
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
Luo, Hanjun
Huang, Zhimu
Chung, Sylvia
Wang, Yiran
Jin, Yingbin
Li, Jialin
Li, Jiang
Li, Xinfeng
Salam, Hanan
Artificial Intelligence
Text-to-image (T2I) systems increasingly rely on upstream prompters, either humans or multimodal large language models (MLLMs), to translate user intent into detailed prompts. Yet current benchmarks fix the prompt and only evaluate T2I models, leaving the prompting proficiency of this upstream component entirely unmeasured. We introduce AtelierEval, the first unified benchmark that quantifies prompting proficiency across 360 expert-crafted tasks. Grounded in a cognitive view, it spans three task categories and instantiates tasks using a taxonomy of real-world challenges, with a dual interface for both humans and MLLMs. To enable scalable and reliable evaluation, we propose AtelierJudge, a skill-based, memory-augmented agentic evaluator. It produces subjective and objective scores for prompt-image pairs, achieving a Spearman correlation of 0.79 with human experts, approaching human performance. Extensive experiments benchmark 8 MLLMs against 48 human users across 4 T2I backends, validate AtelierEval as a robust diagnostic tool, and reveal the superiority of mimicry over planning, advocating for an image-augmented direction for future prompters. Our work is released to support future research.
title AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
topic Artificial Intelligence
url https://arxiv.org/abs/2605.22645