Improve MLLM Benchmark Efficiency through Interview
Fuente:
arXiv
Salvato in:
| Autori principali: | Wen, Farong, Guo, Yijin, Wang, Junying, Xiao, Jiaohao, Zhou, Yingjie, Shen, Ye, Jia, Qi, Li, Chunyi, Zhang, Zicheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
di: Shen, Ye, et al.
Pubblicazione: (2025)
di: Shen, Ye, et al.
Pubblicazione: (2025)
QoNext: Towards Next-generation QoE for Foundation Models
di: Guo, Yijin, et al.
Pubblicazione: (2025)
di: Guo, Yijin, et al.
Pubblicazione: (2025)
Human-Centric Evaluation for Foundation Models
di: Guo, Yijin, et al.
Pubblicazione: (2025)
di: Guo, Yijin, et al.
Pubblicazione: (2025)
The Ever-Evolving Science Exam
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
di: Shen, Ye, et al.
Pubblicazione: (2026)
di: Shen, Ye, et al.
Pubblicazione: (2026)
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
Affordance Benchmark for MLLMs
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
Information Density Principle for MLLM Benchmarks
di: Li, Chunyi, et al.
Pubblicazione: (2025)
di: Li, Chunyi, et al.
Pubblicazione: (2025)
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
di: Zhou, Yingjie, et al.
Pubblicazione: (2024)
di: Zhou, Yingjie, et al.
Pubblicazione: (2024)
SIQA: Toward Reliable Scientific Image Quality Assessment
di: Li, Wenzhe, et al.
Pubblicazione: (2026)
di: Li, Wenzhe, et al.
Pubblicazione: (2026)
Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs
di: Zhu, Xiangyang, et al.
Pubblicazione: (2026)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2026)
User-centric Subjective Leaderboard by Customizable Reward Modeling
di: Jia, Qi, et al.
Pubblicazione: (2025)
di: Jia, Qi, et al.
Pubblicazione: (2025)
Image Quality Assessment for Embodied AI
di: Li, Chunyi, et al.
Pubblicazione: (2025)
di: Li, Chunyi, et al.
Pubblicazione: (2025)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
di: Guo, Haiyun, et al.
Pubblicazione: (2025)
di: Guo, Haiyun, et al.
Pubblicazione: (2025)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
di: Ray, Pretam, et al.
Pubblicazione: (2026)
di: Ray, Pretam, et al.
Pubblicazione: (2026)
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
di: Lu, Xingyu, et al.
Pubblicazione: (2026)
di: Lu, Xingyu, et al.
Pubblicazione: (2026)
STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction
di: Wang, Xiaoxiao, et al.
Pubblicazione: (2026)
di: Wang, Xiaoxiao, et al.
Pubblicazione: (2026)
Redundancy Principles for MLLMs Benchmarks
di: Zhang, Zicheng, et al.
Pubblicazione: (2025)
di: Zhang, Zicheng, et al.
Pubblicazione: (2025)
Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
di: Xiong, Shengwu., et al.
Pubblicazione: (2025)
di: Xiong, Shengwu., et al.
Pubblicazione: (2025)
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
di: Ge, Wentao, et al.
Pubblicazione: (2023)
di: Ge, Wentao, et al.
Pubblicazione: (2023)
Beyond Cosine Similarity: Zero-Initialized Residual Complex Projection for Aspect-Based Sentiment Analysis
di: Wang, Yijin, et al.
Pubblicazione: (2026)
di: Wang, Yijin, et al.
Pubblicazione: (2026)
InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
di: Wang, Xintao, et al.
Pubblicazione: (2023)
di: Wang, Xintao, et al.
Pubblicazione: (2023)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
di: Xie, Wenya, et al.
Pubblicazione: (2025)
di: Xie, Wenya, et al.
Pubblicazione: (2025)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
di: Kil, Jihyung, et al.
Pubblicazione: (2024)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
3DGCQA: A Quality Assessment Database for 3D AI-Generated Contents
di: Zhou, Yingjie, et al.
Pubblicazione: (2024)
di: Zhou, Yingjie, et al.
Pubblicazione: (2024)
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
di: Tan, Chenkun, et al.
Pubblicazione: (2025)
di: Tan, Chenkun, et al.
Pubblicazione: (2025)
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
di: Xiao, Han, et al.
Pubblicazione: (2025)
di: Xiao, Han, et al.
Pubblicazione: (2025)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
di: Yang, Yuwei, et al.
Pubblicazione: (2025)
di: Yang, Yuwei, et al.
Pubblicazione: (2025)
Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning
di: Ni, Yijin, et al.
Pubblicazione: (2025)
di: Ni, Yijin, et al.
Pubblicazione: (2025)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
di: Bei, Yuanchen, et al.
Pubblicazione: (2026)
di: Bei, Yuanchen, et al.
Pubblicazione: (2026)
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
di: Zhou, Yuhao, et al.
Pubblicazione: (2025)
di: Zhou, Yuhao, et al.
Pubblicazione: (2025)
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
di: Jia, Qi, et al.
Pubblicazione: (2025)
di: Jia, Qi, et al.
Pubblicazione: (2025)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
di: Jia, Qi, et al.
Pubblicazione: (2026)
di: Jia, Qi, et al.
Pubblicazione: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
UnifiedMLLM: Enabling Unified Representation for Multi-modal Multi-tasks With Large Language Model
di: Li, Zhaowei, et al.
Pubblicazione: (2024)
di: Li, Zhaowei, et al.
Pubblicazione: (2024)
Comments as Natural Logic Pivots: Improve Code Generation via Comment Perspective
di: Chen, Yijie, et al.
Pubblicazione: (2024)
di: Chen, Yijie, et al.
Pubblicazione: (2024)
WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation
di: Zheng, Zinan, et al.
Pubblicazione: (2026)
di: Zheng, Zinan, et al.
Pubblicazione: (2026)
Judge Anything: MLLM as a Judge Across Any Modality
di: Pu, Shu, et al.
Pubblicazione: (2025)
di: Pu, Shu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
di: Shen, Ye, et al.
Pubblicazione: (2025) -
QoNext: Towards Next-generation QoE for Foundation Models
di: Guo, Yijin, et al.
Pubblicazione: (2025) -
Human-Centric Evaluation for Foundation Models
di: Guo, Yijin, et al.
Pubblicazione: (2025) -
The Ever-Evolving Science Exam
di: Wang, Junying, et al.
Pubblicazione: (2025) -
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
di: Shen, Ye, et al.
Pubblicazione: (2026)