T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Zi-Ao, Lan, Tian, Tu, Rong-Cheng, Liu, Shu-Hang, Huang, Heyan, Wu, Zhijing, Xu, Chen, Mao, Xian-Ling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
by: Tu, Rong-Cheng, et al.
Published: (2024)
by: Tu, Rong-Cheng, et al.
Published: (2024)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
by: Ma, Zi-Ao, et al.
Published: (2024)
by: Ma, Zi-Ao, et al.
Published: (2024)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)
by: Zhang, Guo-Biao, et al.
Published: (2026)
SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection
by: Lu, Yi-Fan, et al.
Published: (2025)
by: Lu, Yi-Fan, et al.
Published: (2025)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
by: Liu, Runheng, et al.
Published: (2026)
by: Liu, Runheng, et al.
Published: (2026)
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models
by: Lu, Yi-Fan, et al.
Published: (2024)
by: Lu, Yi-Fan, et al.
Published: (2024)
Mix-Initiative Response Generation with Dynamic Prefix Tuning
by: Nie, Yuxiang, et al.
Published: (2024)
by: Nie, Yuxiang, et al.
Published: (2024)
Distribution-Consistency-Guided Multi-modal Hashing
by: Liu, Jin-Yu, et al.
Published: (2024)
by: Liu, Jin-Yu, et al.
Published: (2024)
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
by: Chen, Kaijie, et al.
Published: (2025)
by: Chen, Kaijie, et al.
Published: (2025)
A Distributed Collaborative Retrieval Framework Excelling in All Queries and Corpora based on Zero-shot Rank-Oriented Automatic Evaluation
by: Che, Tian-Yi, et al.
Published: (2024)
by: Che, Tian-Yi, et al.
Published: (2024)
Training-free Truthfulness Detection via Value Vectors in LLMs
by: Liu, Runheng, et al.
Published: (2025)
by: Liu, Runheng, et al.
Published: (2025)
T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2025)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2025)
DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models
by: Wang, Juntong, et al.
Published: (2026)
by: Wang, Juntong, et al.
Published: (2026)
Training Language Models to Critique With Multi-agent Feedback
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
EXCEEDS: Extracting Complex Events via Nugget-based Grid Modeling in Scientific Domain
by: Lu, Yi-Fan, et al.
Published: (2024)
by: Lu, Yi-Fan, et al.
Published: (2024)
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
by: Sun, Kaiyue, et al.
Published: (2025)
by: Sun, Kaiyue, et al.
Published: (2025)
ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation
by: Sun, Rui-Qing, et al.
Published: (2025)
by: Sun, Rui-Qing, et al.
Published: (2025)
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
by: Liu, Yilun, et al.
Published: (2026)
by: Liu, Yilun, et al.
Published: (2026)
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
by: Tian, Wei, et al.
Published: (2026)
by: Tian, Wei, et al.
Published: (2026)
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation
by: Tu, Quan, et al.
Published: (2024)
by: Tu, Quan, et al.
Published: (2024)
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
by: Zhou, Youchao, et al.
Published: (2024)
by: Zhou, Youchao, et al.
Published: (2024)
FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
by: Liu, Runheng, et al.
Published: (2024)
by: Liu, Runheng, et al.
Published: (2024)
StrucText-Eval: Evaluating Large Language Model's Reasoning Ability in Structure-Rich Text
by: Gu, Zhouhong, et al.
Published: (2024)
by: Gu, Zhouhong, et al.
Published: (2024)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
by: Xu, Yumo, et al.
Published: (2025)
by: Xu, Yumo, et al.
Published: (2025)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
by: Kamath, Amita, et al.
Published: (2025)
by: Kamath, Amita, et al.
Published: (2025)
AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters
by: Luo, Hanjun, et al.
Published: (2026)
by: Luo, Hanjun, et al.
Published: (2026)
ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning
by: Chen, Honghua, et al.
Published: (2026)
by: Chen, Honghua, et al.
Published: (2026)
Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation
by: Yao, Jiashu, et al.
Published: (2026)
by: Yao, Jiashu, et al.
Published: (2026)
Efficient and Robust Video Defense Framework against 3D-field Personalized Talking Face
by: Sun, Rui-qing, et al.
Published: (2025)
by: Sun, Rui-qing, et al.
Published: (2025)
EvalQReason: A Framework for Step-Level Reasoning Evaluation in Large Language Models
by: Freja, Shaima Ahmad, et al.
Published: (2026)
by: Freja, Shaima Ahmad, et al.
Published: (2026)
SparseEval: Efficient Evaluation of Large Language Models by Sparse Optimization
by: Zhang, Taolin, et al.
Published: (2026)
by: Zhang, Taolin, et al.
Published: (2026)
R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
by: Cheng, Aijia, et al.
Published: (2026)
by: Cheng, Aijia, et al.
Published: (2026)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
MMWOZ: Building Multimodal Agent for Task-oriented Dialogue
by: Yang, Pu-Hai, et al.
Published: (2025)
by: Yang, Pu-Hai, et al.
Published: (2025)
All‐Optical Imaging Using Perovskite Nanocrystals Based on Spectro‐Spatial Correlation
by: Xuhong Wang, et al.
Published: (2024)
by: Xuhong Wang, et al.
Published: (2024)
TextShield-R1: Reinforced Reasoning for Tampered Text Detection
by: Qu, Chenfan, et al.
Published: (2026)
by: Qu, Chenfan, et al.
Published: (2026)
Similar Items
-
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
by: Tu, Rong-Cheng, et al.
Published: (2024) -
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
by: Ma, Zi-Ao, et al.
Published: (2024) -
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
by: Lan, Tian, et al.
Published: (2025) -
CriticEval: Evaluating Large Language Model as Critic
by: Lan, Tian, et al.
Published: (2024) -
DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey
by: Zhang, Guo-Biao, et al.
Published: (2026)