Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koo, Jaywon, Hernandez, Jefferson, Haji-Ali, Moayed, Yang, Ziyan, Ordonez, Vicente
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909754220085248
author Koo, Jaywon
Hernandez, Jefferson
Haji-Ali, Moayed
Yang, Ziyan
Ordonez, Vicente
author_facet Koo, Jaywon
Hernandez, Jefferson
Haji-Ali, Moayed
Yang, Ziyan
Ordonez, Vicente
contents Evaluating text-to-image and text-to-video models is challenging due to a fundamental disconnect: established metrics fail to jointly measure visual quality and semantic alignment with text, leading to a poor correlation with human judgments. To address this critical issue, we propose cFreD, a general metric based on a Conditional Fréchet Distance that unifies the assessment of visual fidelity and text-prompt consistency into a single score. Existing metrics such as Fréchet Inception Distance (FID) capture image quality but ignore text conditioning while alignment scores such as CLIPScore are insensitive to visual quality. Furthermore, learned preference models require constant retraining and are unlikely to generalize to novel architectures or out-of-distribution prompts. Through extensive experiments across multiple recently proposed text-to-image models and diverse prompt datasets, cFreD exhibits a higher correlation with human judgments compared to statistical metrics , including metrics trained with human preferences. Our findings validate cFreD as a robust, future-proof metric for the systematic evaluation of text conditioned models, standardizing benchmarking in this rapidly evolving field. We release our evaluation toolkit and benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
Koo, Jaywon
Hernandez, Jefferson
Haji-Ali, Moayed
Yang, Ziyan
Ordonez, Vicente
Computer Vision and Pattern Recognition
Evaluating text-to-image and text-to-video models is challenging due to a fundamental disconnect: established metrics fail to jointly measure visual quality and semantic alignment with text, leading to a poor correlation with human judgments. To address this critical issue, we propose cFreD, a general metric based on a Conditional Fréchet Distance that unifies the assessment of visual fidelity and text-prompt consistency into a single score. Existing metrics such as Fréchet Inception Distance (FID) capture image quality but ignore text conditioning while alignment scores such as CLIPScore are insensitive to visual quality. Furthermore, learned preference models require constant retraining and are unlikely to generalize to novel architectures or out-of-distribution prompts. Through extensive experiments across multiple recently proposed text-to-image models and diverse prompt datasets, cFreD exhibits a higher correlation with human judgments compared to statistical metrics , including metrics trained with human preferences. Our findings validate cFreD as a robust, future-proof metric for the systematic evaluation of text conditioned models, standardizing benchmarking in this rapidly evolving field. We release our evaluation toolkit and benchmark.
title Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.21721