What makes a good metric? Evaluating automatic metrics for text-to-image consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Ross, Candace, Hall, Melissa, Soriano, Adriana Romero, Williams, Adina |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023)
by: Hall, Melissa, et al.
Published: (2023)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
by: Mañas, Oscar, et al.
Published: (2024)
by: Mañas, Oscar, et al.
Published: (2024)
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
Evaluating the Evaluators: Are readability metrics good measures of readability?
by: Cachola, Isabel, et al.
Published: (2025)
by: Cachola, Isabel, et al.
Published: (2025)
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Decomposed evaluations of geographic disparities in text-to-image models
by: Sureddy, Abhishek, et al.
Published: (2024)
by: Sureddy, Abhishek, et al.
Published: (2024)
Serialized EHR make for good text representations
by: Chou, Zhirong, et al.
Published: (2025)
by: Chou, Zhirong, et al.
Published: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
The statistical advantage of automatic NLG metrics at the system level
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
Domain Regeneration: How well do LLMs match syntactic properties of text domains?
by: Ju, Da, et al.
Published: (2025)
by: Ju, Da, et al.
Published: (2025)
Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages
by: Ovalle, Anaelia, et al.
Published: (2025)
by: Ovalle, Anaelia, et al.
Published: (2025)
How do we measure privacy in text? A survey of text anonymization metrics
by: Ren, Yaxuan, et al.
Published: (2025)
by: Ren, Yaxuan, et al.
Published: (2025)
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
by: Teotia, Revant, et al.
Published: (2025)
by: Teotia, Revant, et al.
Published: (2025)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025)
by: Ross, Candace, et al.
Published: (2025)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
by: Hemmat, Reyhane Askari, et al.
Published: (2024)
What do the metrics mean? A critical analysis of the use of Automated Evaluation Metrics in Interpreting
by: Downie, Jonathan, et al.
Published: (2026)
by: Downie, Jonathan, et al.
Published: (2026)
[Call for Papers] The 2nd BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
by: Gao, Peiheng, et al.
Published: (2025)
by: Gao, Peiheng, et al.
Published: (2025)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
by: Ichmoukhamedov, Timour, et al.
Published: (2024)
by: Ichmoukhamedov, Timour, et al.
Published: (2024)
*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
by: Lemesle, Quentin, et al.
Published: (2026)
by: Lemesle, Quentin, et al.
Published: (2026)
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
by: Hu, Michael Y., et al.
Published: (2024)
by: Hu, Michael Y., et al.
Published: (2024)
Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics
by: Kocbek, Primoz, et al.
Published: (2025)
by: Kocbek, Primoz, et al.
Published: (2025)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
by: Goto, Takumi, et al.
Published: (2025)
by: Goto, Takumi, et al.
Published: (2025)
Are Female Carpenters like Blue Bananas? A Corpus Investigation of Occupation Gender Typicality
by: Ju, Da, et al.
Published: (2024)
by: Ju, Da, et al.
Published: (2024)
The illusion of a perfect metric: Why evaluating AI's words is harder than it looks
by: Oliva, Maria Paz, et al.
Published: (2025)
by: Oliva, Maria Paz, et al.
Published: (2025)
Classifying populist language in American presidential and governor speeches using automatic text analysis
by: van der Veen, Olaf, et al.
Published: (2024)
by: van der Veen, Olaf, et al.
Published: (2024)
Measuring text summarization factuality using atomic facts entailment metrics in the context of retrieval augmented generation
by: Kriman, N. E.
Published: (2024)
by: Kriman, N. E.
Published: (2024)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
by: Kerkouri, Mohamed Amine, et al.
Published: (2026)
What makes an entity salient in discourse?
by: Zeldes, Amir, et al.
Published: (2025)
by: Zeldes, Amir, et al.
Published: (2025)
Quantifying consistency and accuracy of Latent Dirichlet Allocation
by: Magsarjav, Saranzaya, et al.
Published: (2025)
by: Magsarjav, Saranzaya, et al.
Published: (2025)
BabyLM Turns 3: Call for papers for the 2025 BabyLM workshop
by: Charpentier, Lucas, et al.
Published: (2025)
by: Charpentier, Lucas, et al.
Published: (2025)
Scalable and consistent few-shot classification of survey responses using text embeddings
by: Mjaaland, Jonas Timmann, et al.
Published: (2025)
by: Mjaaland, Jonas Timmann, et al.
Published: (2025)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Domain-specific or Uncertainty-aware models: Does it really make a difference for biomedical text classification?
by: Sinha, Aman, et al.
Published: (2024)
by: Sinha, Aman, et al.
Published: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
by: Cunha, Washington, et al.
Published: (2025)
by: Cunha, Washington, et al.
Published: (2025)
Developing an AI framework to automatically detect shared decision-making in patient-doctor conversations
by: Ponce-Ponte, Oscar J., et al.
Published: (2025)
by: Ponce-Ponte, Oscar J., et al.
Published: (2025)
A review of faithfulness metrics for hallucination assessment in Large Language Models
by: Malin, Ben, et al.
Published: (2024)
by: Malin, Ben, et al.
Published: (2024)
Visual question answering based evaluation metrics for text-to-image generation
by: Miyamoto, Mizuki, et al.
Published: (2024)
by: Miyamoto, Mizuki, et al.
Published: (2024)
Similar Items
-
DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity
by: Hall, Melissa, et al.
Published: (2023) -
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024) -
Improving Text-to-Image Consistency via Automatic Prompt Optimization
by: Mañas, Oscar, et al.
Published: (2024) -
Multi-Modal Language Models as Text-to-Image Model Evaluators
by: Chen, Jiahui, et al.
Published: (2025) -
Evaluating the Evaluators: Are readability metrics good measures of readability?
by: Cachola, Isabel, et al.
Published: (2025)