SedarEval: Automated Evaluation using Self-Adaptive Rubrics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fan, Zhiyuan, Wang, Weinong, Wu, Xing, Zhang, Debing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
von: Chen, Jiahua, et al.
Veröffentlicht: (2026)
Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning
von: Zhang, Qin, et al.
Veröffentlicht: (2026)
von: Zhang, Qin, et al.
Veröffentlicht: (2026)
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
von: Ni, Minheng, et al.
Veröffentlicht: (2025)
TactileEval: A Step Towards Automated Fine-Grained Evaluation and Editing of Tactile Graphics
von: Khan, Adnan, et al.
Veröffentlicht: (2026)
von: Khan, Adnan, et al.
Veröffentlicht: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
von: Wang, Pengfei, et al.
Veröffentlicht: (2025)
von: Wang, Pengfei, et al.
Veröffentlicht: (2025)
EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
von: Liu, Yaofang, et al.
Veröffentlicht: (2023)
von: Liu, Yaofang, et al.
Veröffentlicht: (2023)
VideoGen-Eval: Agent-based System for Video Generation Evaluation
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
von: Yang, Yuhang, et al.
Veröffentlicht: (2025)
Visual Preference Optimization with Rubric Rewards
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Robust Rubric Rewards
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
von: Yu, Ya-Qi, et al.
Veröffentlicht: (2026)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
von: Ma, Wentao, et al.
Veröffentlicht: (2025)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2025)
PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior
von: Wu, Junda, et al.
Veröffentlicht: (2025)
von: Wu, Junda, et al.
Veröffentlicht: (2025)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
von: Li, Xinhao, et al.
Veröffentlicht: (2024)
Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation
von: Gupta, Arnav, et al.
Veröffentlicht: (2025)
von: Gupta, Arnav, et al.
Veröffentlicht: (2025)
DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models
von: Wang, Juntong, et al.
Veröffentlicht: (2026)
von: Wang, Juntong, et al.
Veröffentlicht: (2026)
HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation
von: Zhang, Bingzi, et al.
Veröffentlicht: (2026)
von: Zhang, Bingzi, et al.
Veröffentlicht: (2026)
DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
von: Babu, Nithin C., et al.
Veröffentlicht: (2025)
von: Babu, Nithin C., et al.
Veröffentlicht: (2025)
EvalGIM: A Library for Evaluating Generative Image Models
von: Hall, Melissa, et al.
Veröffentlicht: (2024)
von: Hall, Melissa, et al.
Veröffentlicht: (2024)
CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
von: Wang, Chonghuinan, et al.
Veröffentlicht: (2026)
von: Wang, Chonghuinan, et al.
Veröffentlicht: (2026)
RICA2: Rubric-Informed, Calibrated Assessment of Actions
von: Majeedi, Abrar, et al.
Veröffentlicht: (2024)
von: Majeedi, Abrar, et al.
Veröffentlicht: (2024)
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
von: yunhan, Li, et al.
Veröffentlicht: (2025)
von: yunhan, Li, et al.
Veröffentlicht: (2025)
PDSE: A Multiple Lesion Detector for CT Images using PANet and Deformable Squeeze-and-Excitation Block
von: Fan, Di, et al.
Veröffentlicht: (2025)
von: Fan, Di, et al.
Veröffentlicht: (2025)
HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
von: Wang, Zichuan, et al.
Veröffentlicht: (2025)
von: Wang, Zichuan, et al.
Veröffentlicht: (2025)
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
von: Chen, Kai, et al.
Veröffentlicht: (2024)
von: Chen, Kai, et al.
Veröffentlicht: (2024)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
Flow-Anything: Learning Real-World Optical Flow Estimation from Large-Scale Single-view Images
von: Liang, Yingping, et al.
Veröffentlicht: (2025)
von: Liang, Yingping, et al.
Veröffentlicht: (2025)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
von: Zhao, Lin, et al.
Veröffentlicht: (2024)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
NucEval: A Robust Evaluation Framework for Nuclear Instance Segmentation
von: Mahbod, Amirreza, et al.
Veröffentlicht: (2026)
von: Mahbod, Amirreza, et al.
Veröffentlicht: (2026)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
von: Chen, Xin-Sheng, et al.
Veröffentlicht: (2026)
von: Chen, Xin-Sheng, et al.
Veröffentlicht: (2026)
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhou, Yingjie, et al.
Veröffentlicht: (2025)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
von: Tang, Guowei, et al.
Veröffentlicht: (2026)
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning
von: Chen, Jiahua, et al.
Veröffentlicht: (2026) -
Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning
von: Zhang, Qin, et al.
Veröffentlicht: (2026) -
TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation
von: Ni, Minheng, et al.
Veröffentlicht: (2025) -
TactileEval: A Step Towards Automated Fine-Grained Evaluation and Editing of Tactile Graphics
von: Khan, Adnan, et al.
Veröffentlicht: (2026) -
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)