Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Tu, Rong-Cheng, Ma, Zi-Ao, Lan, Tian, Zhao, Yuehao, Huang, Heyan, Mao, Xian-Ling |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
by: Ma, Zi-Ao, et al.
Published: (2025)
by: Ma, Zi-Ao, et al.
Published: (2025)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
by: Ma, Zi-Ao, et al.
Published: (2024)
by: Ma, Zi-Ao, et al.
Published: (2024)
SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection
by: Lu, Yi-Fan, et al.
Published: (2025)
by: Lu, Yi-Fan, et al.
Published: (2025)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
Training Language Models to Critique With Multi-agent Feedback
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models
by: Lu, Yi-Fan, et al.
Published: (2024)
by: Lu, Yi-Fan, et al.
Published: (2024)
Mix-Initiative Response Generation with Dynamic Prefix Tuning
by: Nie, Yuxiang, et al.
Published: (2024)
by: Nie, Yuxiang, et al.
Published: (2024)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
MMWOZ: Building Multimodal Agent for Task-oriented Dialogue
by: Yang, Pu-Hai, et al.
Published: (2025)
by: Yang, Pu-Hai, et al.
Published: (2025)
Distribution-Consistency-Guided Multi-modal Hashing
by: Liu, Jin-Yu, et al.
Published: (2024)
by: Liu, Jin-Yu, et al.
Published: (2024)
EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain
by: Lu, Yi-Fan, et al.
Published: (2024)
by: Lu, Yi-Fan, et al.
Published: (2024)
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability
by: Hu, Xinyu, et al.
Published: (2025)
by: Hu, Xinyu, et al.
Published: (2025)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
by: Tian, Yanzhi, et al.
Published: (2026)
by: Tian, Yanzhi, et al.
Published: (2026)
Dataset Distillation by Automatic Training Trajectories
by: Liu, Dai, et al.
Published: (2024)
by: Liu, Dai, et al.
Published: (2024)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
by: Zhao, Shihao, et al.
Published: (2025)
by: Zhao, Shihao, et al.
Published: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
by: Zhou, Xiaofeng, et al.
Published: (2025)
by: Zhou, Xiaofeng, et al.
Published: (2025)
Cross-Architecture Distillation Made Simple with Redundancy Suppression
by: Zhang, Weijia, et al.
Published: (2025)
by: Zhang, Weijia, et al.
Published: (2025)
A Distributed Collaborative Retrieval Framework Excelling in All Queries and Corpora based on Zero-shot Rank-Oriented Automatic Evaluation
by: Che, Tian-Yi, et al.
Published: (2024)
by: Che, Tian-Yi, et al.
Published: (2024)
TextSculptor: Training and Benchmarking Scene Text Editing
by: Lin, Yiheng, et al.
Published: (2026)
by: Lin, Yiheng, et al.
Published: (2026)
Automatic Extraction of Metaphoric Analogies from Literary Texts: Task Formulation, Dataset Construction, and Evaluation
by: Boisson, Joanne, et al.
Published: (2024)
by: Boisson, Joanne, et al.
Published: (2024)
Training-free Truthfulness Detection via Value Vectors in LLMs
by: Liu, Runheng, et al.
Published: (2025)
by: Liu, Runheng, et al.
Published: (2025)
Training Task Experts through Retrieval Based Distillation
by: Ge, Jiaxin, et al.
Published: (2024)
by: Ge, Jiaxin, et al.
Published: (2024)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
by: Tu, Quan, et al.
Published: (2024)
by: Tu, Quan, et al.
Published: (2024)
The COTe score: A decomposable framework for evaluating Document Layout Analysis models
by: Bourne, Jonathan, et al.
Published: (2026)
by: Bourne, Jonathan, et al.
Published: (2026)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models
by: Han, Simeng, et al.
Published: (2025)
by: Han, Simeng, et al.
Published: (2025)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
by: Liu, Runheng, et al.
Published: (2026)
by: Liu, Runheng, et al.
Published: (2026)
Retrieval Or Holistic Understanding? Dolce: Differentiate Our Long Context Evaluation Tasks
by: Yang, Zi
Published: (2024)
by: Yang, Zi
Published: (2024)
SLVMEval: Synthetic Meta Evaluation Benchmark for Text-to-Long Video Generation
by: Matsuda, Ryosuke, et al.
Published: (2026)
by: Matsuda, Ryosuke, et al.
Published: (2026)
ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
by: Lan, Qizhen, et al.
Published: (2025)
by: Lan, Qizhen, et al.
Published: (2025)
Efficient and Robust Video Defense Framework against 3D-field Personalized Talking Face
by: Sun, Rui-qing, et al.
Published: (2025)
by: Sun, Rui-qing, et al.
Published: (2025)
MetaIE: Distilling a Meta Model from LLM for All Kinds of Information Extraction Tasks
by: Peng, Letian, et al.
Published: (2024)
by: Peng, Letian, et al.
Published: (2024)
LLMs Judge Themselves: A Game-Theoretic Framework for Human-Aligned Evaluation
by: Yang, Gao, et al.
Published: (2025)
by: Yang, Gao, et al.
Published: (2025)
CLoCKDistill: Consistent Location-and-Context-aware Knowledge Distillation for DETRs
by: Lan, Qizhen, et al.
Published: (2025)
by: Lan, Qizhen, et al.
Published: (2025)
Similar Items
-
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025) -
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
by: Ma, Zi-Ao, et al.
Published: (2025) -
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
by: Ma, Zi-Ao, et al.
Published: (2024) -
SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection
by: Lu, Yi-Fan, et al.
Published: (2025) -
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)