Constantly Improving Image Models Need Constantly Improving Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Jiaxin, Luo, Grace, Lee, Heekyung, Malpani, Nishant, Lian, Long, Wang, XuDong, Holynski, Aleksander, Darrell, Trevor, Min, Sewon, Chan, David M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909852108849152
author Ge, Jiaxin
Luo, Grace
Lee, Heekyung
Malpani, Nishant
Lian, Long
Wang, XuDong
Holynski, Aleksander
Darrell, Trevor
Min, Sewon
Chan, David M.
author_facet Ge, Jiaxin
Luo, Grace
Lee, Heekyung
Malpani, Nishant
Lian, Long
Wang, XuDong
Holynski, Aleksander
Darrell, Trevor
Min, Sewon
Chan, David M.
contents Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture these emerging use cases, leaving a gap between community perceptions of progress and formal evaluation. To address this, we present ECHO, a framework for constructing benchmarks directly from real-world evidence of model use: social media posts that showcase novel prompts and qualitative user judgments. Applying this framework to GPT-4o Image Gen, we construct a dataset of over 31,000 prompts curated from such posts. Our analysis shows that ECHO (1) discovers creative and complex tasks absent from existing benchmarks, such as re-rendering product labels across languages or generating receipts with specified totals, (2) more clearly distinguishes state-of-the-art models from alternatives, and (3) surfaces community feedback that we use to inform the design of metrics for model quality (e.g., measuring observed shifts in color, identity, and structure). Our website is at https://echo-bench.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Constantly Improving Image Models Need Constantly Improving Benchmarks
Ge, Jiaxin
Luo, Grace
Lee, Heekyung
Malpani, Nishant
Lian, Long
Wang, XuDong
Holynski, Aleksander
Darrell, Trevor
Min, Sewon
Chan, David M.
Computer Vision and Pattern Recognition
Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture these emerging use cases, leaving a gap between community perceptions of progress and formal evaluation. To address this, we present ECHO, a framework for constructing benchmarks directly from real-world evidence of model use: social media posts that showcase novel prompts and qualitative user judgments. Applying this framework to GPT-4o Image Gen, we construct a dataset of over 31,000 prompts curated from such posts. Our analysis shows that ECHO (1) discovers creative and complex tasks absent from existing benchmarks, such as re-rendering product labels across languages or generating receipts with specified totals, (2) more clearly distinguishes state-of-the-art models from alternatives, and (3) surfaces community feedback that we use to inform the design of metrics for model quality (e.g., measuring observed shifts in color, identity, and structure). Our website is at https://echo-bench.github.io.
title Constantly Improving Image Models Need Constantly Improving Benchmarks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.15021