Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Huixuan, Wan, Xiaojun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916787742834688
author Zhang, Huixuan
Wan, Xiaojun
author_facet Zhang, Huixuan
Wan, Xiaojun
contents Text-to-image models often struggle to generate images that precisely match textual prompts. Prior research has extensively studied the evaluation of image-text alignment in text-to-image generation. However, existing evaluations primarily focus on agreement with human assessments, neglecting other critical properties of a trustworthy evaluation framework. In this work, we first identify two key aspects that a reliable evaluation should address. We then empirically demonstrate that current mainstream evaluation frameworks fail to fully satisfy these properties across a diverse range of metrics and models. Finally, we propose recommendations for improving image-text alignment evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08480
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
Zhang, Huixuan
Wan, Xiaojun
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Text-to-image models often struggle to generate images that precisely match textual prompts. Prior research has extensively studied the evaluation of image-text alignment in text-to-image generation. However, existing evaluations primarily focus on agreement with human assessments, neglecting other critical properties of a trustworthy evaluation framework. In this work, we first identify two key aspects that a reliable evaluation should address. We then empirically demonstrate that current mainstream evaluation frameworks fail to fully satisfy these properties across a diverse range of metrics and models. Finally, we propose recommendations for improving image-text alignment evaluation.
title Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08480