Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Botian, Li, Lei, Li, Xiaonan, Li, Zhaowei, Feng, Xiachong, Kong, Lingpeng, Liu, Qi, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reasoning Does Not Necessarily Improve Role-Playing Ability
by: Feng, Xiachong, et al.
Published: (2025)
by: Feng, Xiachong, et al.
Published: (2025)
The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models
by: Qin, Chonghan, et al.
Published: (2026)
by: Qin, Chonghan, et al.
Published: (2026)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
by: Wang, Haochuan, et al.
Published: (2024)
by: Wang, Haochuan, et al.
Published: (2024)
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
by: Sun, Qiushi, et al.
Published: (2023)
by: Sun, Qiushi, et al.
Published: (2023)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
by: Zong, Yi, et al.
Published: (2024)
by: Zong, Yi, et al.
Published: (2024)
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
by: Zhou, Yitong, et al.
Published: (2025)
by: Zhou, Yitong, et al.
Published: (2025)
A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios
by: Feng, Xiachong, et al.
Published: (2024)
by: Feng, Xiachong, et al.
Published: (2024)
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
by: Zhang, Xiaotian, et al.
Published: (2023)
by: Zhang, Xiaotian, et al.
Published: (2023)
LLatrieval: LLM-Verified Retrieval for Verifiable Generation
by: Li, Xiaonan, et al.
Published: (2023)
by: Li, Xiaonan, et al.
Published: (2023)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
by: Yang, Chenchen, et al.
Published: (2026)
by: Yang, Chenchen, et al.
Published: (2026)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
by: Fei, Zhaoye, et al.
Published: (2025)
by: Fei, Zhaoye, et al.
Published: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
by: Sun, Qiushi, et al.
Published: (2025)
by: Sun, Qiushi, et al.
Published: (2025)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
Flames: Benchmarking Value Alignment of LLMs in Chinese
by: Huang, Kexin, et al.
Published: (2023)
by: Huang, Kexin, et al.
Published: (2023)
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
by: Prange, Jakob, et al.
Published: (2021)
by: Prange, Jakob, et al.
Published: (2021)
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)
by: Li, Shimin, et al.
Published: (2024)
CURE: A Multimodal Benchmark for Clinical Understanding and Retrieval Evaluation
by: Gu, Yannian, et al.
Published: (2026)
by: Gu, Yannian, et al.
Published: (2026)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
by: Xu, Ningyu, et al.
Published: (2026)
by: Xu, Ningyu, et al.
Published: (2026)
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
by: Jiang, Zhenhui, et al.
Published: (2024)
by: Jiang, Zhenhui, et al.
Published: (2024)
Proxy Compression for Language Modeling
by: Zheng, Lin, et al.
Published: (2026)
by: Zheng, Lin, et al.
Published: (2026)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
by: Yan, Weixiang, et al.
Published: (2023)
by: Yan, Weixiang, et al.
Published: (2023)
Identifying Semantic Induction Heads to Understand In-Context Learning
by: Ren, Jie, et al.
Published: (2024)
by: Ren, Jie, et al.
Published: (2024)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
by: Ye, Yangfan, et al.
Published: (2024)
by: Ye, Yangfan, et al.
Published: (2024)
Nondeterministic Polynomial-time Problem Challenge: An Ever-Scaling Reasoning Benchmark for LLMs
by: Yang, Chang, et al.
Published: (2025)
by: Yang, Chang, et al.
Published: (2025)
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Multimodal Table Understanding
by: Zheng, Mingyu, et al.
Published: (2024)
by: Zheng, Mingyu, et al.
Published: (2024)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Scaling Reasoning without Attention
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating
by: Deng, Chao, et al.
Published: (2024)
by: Deng, Chao, et al.
Published: (2024)
Causal Understanding by LLMs: The Role of Uncertainty
by: Lithgow-Serrano, Oscar, et al.
Published: (2025)
by: Lithgow-Serrano, Oscar, et al.
Published: (2025)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
Similar Items
-
Reasoning Does Not Necessarily Improve Role-Playing Ability
by: Feng, Xiachong, et al.
Published: (2025) -
The Granularity Axis: A Micro-to-Macro Latent Direction for Social Roles in Language Models
by: Qin, Chonghan, et al.
Published: (2026) -
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025) -
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
by: Wang, Haochuan, et al.
Published: (2024) -
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
by: Sun, Qiushi, et al.
Published: (2023)