GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ruihang, Qu, Leigang, Zhang, Jingxu, Gui, Dongnan, Xu, Mengde, Zhang, Xiaosong, Hu, Han, Wang, Wenjie, Wang, Jiaqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comment on: “Obesity increases the risk of major wound complications following pelvic resection for bone sarcoma”
by: Mengde Qu, et al.
Published: (2024)
by: Mengde Qu, et al.
Published: (2024)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Can the stress be managed? Stress mindset as a mitigating factor in the influence of job demands on burnout
by: Yaoying Zhou, et al.
Published: (2024)
by: Yaoying Zhou, et al.
Published: (2024)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
AlignedGen: Aligning Style Across Generated Images
by: Zhang, Jiexuan, et al.
Published: (2025)
by: Zhang, Jiexuan, et al.
Published: (2025)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
by: Yang, Jialiang, et al.
Published: (2026)
by: Yang, Jialiang, et al.
Published: (2026)
MPO: An Efficient Post-Processing Framework for Mixing Diverse Preference Alignment
by: Wang, Tianze, et al.
Published: (2025)
by: Wang, Tianze, et al.
Published: (2025)
Illusions in Humans and AI: How Visual Perception Aligns and Diverges
by: Yang, Jianyi, et al.
Published: (2025)
by: Yang, Jianyi, et al.
Published: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
The AI Hippocampus: How Far are We From Human Memory?
by: Jia, Zixia, et al.
Published: (2026)
by: Jia, Zixia, et al.
Published: (2026)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Dance Any Beat: Blending Beats with Visuals in Dance Video Generation
by: Wang, Xuanchen, et al.
Published: (2024)
by: Wang, Xuanchen, et al.
Published: (2024)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Beauty in the Eye of AI: Aligning LLMs and Vision Models with Human Aesthetics in Network Visualization
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
Aligning Language Models with Human Preferences via a Bayesian Approach
by: Wang, Jiashuo, et al.
Published: (2023)
by: Wang, Jiashuo, et al.
Published: (2023)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
by: Wang, Dianyi, et al.
Published: (2026)
by: Wang, Dianyi, et al.
Published: (2026)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
by: Fang, Qingkai, et al.
Published: (2024)
by: Fang, Qingkai, et al.
Published: (2024)
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
by: Huang, Zhipeng, et al.
Published: (2025)
by: Huang, Zhipeng, et al.
Published: (2025)
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
by: Wang, An-Lan, et al.
Published: (2025)
by: Wang, An-Lan, et al.
Published: (2025)
Enhancing Advanced Visual Reasoning Ability of Large Language Models
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
GenLens: A Systematic Evaluation of Visual GenAI Model Outputs
by: Lin, Tica, et al.
Published: (2024)
by: Lin, Tica, et al.
Published: (2024)
Universal gradient estimates for solutions of $Δ_{p,f}u+au^σ\ln u=0$ on complete Riemannian manifolds
by: Liu, Jingxu, et al.
Published: (2026)
by: Liu, Jingxu, et al.
Published: (2026)
How Far Can We Go with Practical Function-Level Program Repair?
by: Xiang, Jiahong, et al.
Published: (2024)
by: Xiang, Jiahong, et al.
Published: (2024)
GenML: A Python Library to Generate the Mittag-Leffler Correlated Noise
by: Qu, Xiang, et al.
Published: (2024)
by: Qu, Xiang, et al.
Published: (2024)
NeuroGen: Neural Network Parameter Generation via Large Language Models
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Position: AI Evaluation Should Learn from How We Test Humans
by: Zhuang, Yan, et al.
Published: (2023)
by: Zhuang, Yan, et al.
Published: (2023)
Multi-modal Data Binding for Survival Analysis Modeling with Incomplete Data and Annotations
by: Qu, Linhao, et al.
Published: (2024)
by: Qu, Linhao, et al.
Published: (2024)
A Scoping Review on Goals of Care Discussions in Surgery: How Are We Doing and How Can We Do Better?
by: Amanda Mac, et al.
Published: (2025)
by: Amanda Mac, et al.
Published: (2025)
Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
by: Xu, Ruihang, et al.
Published: (2025)
by: Xu, Ruihang, et al.
Published: (2025)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
Debias Can be Unreliable: Mitigating Bias Issue in Evaluating Debiasing Recommendation
by: Wang, Chengbing, et al.
Published: (2024)
by: Wang, Chengbing, et al.
Published: (2024)
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
by: Ge, Yate, et al.
Published: (2025)
by: Ge, Yate, et al.
Published: (2025)
Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study
by: Shao, Zekai, et al.
Published: (2025)
by: Shao, Zekai, et al.
Published: (2025)
Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
by: Qi, Minfeng, et al.
Published: (2025)
by: Qi, Minfeng, et al.
Published: (2025)
Similar Items
-
Comment on: “Obesity increases the risk of major wound complications following pelvic resection for bone sarcoma”
by: Mengde Qu, et al.
Published: (2024) -
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024) -
Discriminative Probing and Tuning for Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024) -
Can the stress be managed? Stress mindset as a mitigating factor in the influence of job demands on burnout
by: Yaoying Zhou, et al.
Published: (2024) -
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
by: Qu, Leigang, et al.
Published: (2025)