Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kasaei, Seyed Amir, Rohban, Mohammad Hossein
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918191535489024
author Kasaei, Seyed Amir
Rohban, Mohammad Hossein
author_facet Kasaei, Seyed Amir
Rohban, Mohammad Hossein
contents In language and vision-language models, hallucination is broadly understood as content generated from a model's prior knowledge or biases rather than from the given input. While this phenomenon has been studied in those domains, it has not been clearly framed for text-to-image (T2I) generative models. Existing evaluations mainly focus on alignment, checking whether prompt-specified elements appear, but overlook what the model generates beyond the prompt. We argue for defining hallucination in T2I as bias-driven deviations and propose a taxonomy with three categories: attribute, relation, and object hallucinations. This framing introduces an upper bound for evaluation and surfaces hidden biases, providing a foundation for richer assessment of T2I models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21257
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
Kasaei, Seyed Amir
Rohban, Mohammad Hossein
Computer Vision and Pattern Recognition
Computation and Language
In language and vision-language models, hallucination is broadly understood as content generated from a model's prior knowledge or biases rather than from the given input. While this phenomenon has been studied in those domains, it has not been clearly framed for text-to-image (T2I) generative models. Existing evaluations mainly focus on alignment, checking whether prompt-specified elements appear, but overlook what the model generates beyond the prompt. We argue for defining hallucination in T2I as bias-driven deviations and propose a taxonomy with three categories: attribute, relation, and object hallucinations. This framing introduces an upper bound for evaluation and surfaces hidden biases, providing a foundation for richer assessment of T2I models.
title Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.21257