MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Fuente:
arXiv
Saved in:
| Main Authors: | Yue, Xiang, Ni, Yuansheng, Zhang, Kai, Zheng, Tianyu, Liu, Ruoqi, Zhang, Ge, Stevens, Samuel, Jiang, Dongfu, Ren, Weiming, Sun, Yuxuan, Wei, Cong, Yu, Botao, Yuan, Ruibin, Sun, Renliang, Yin, Ming, Zheng, Boyuan, Yang, Zhenzhu, Liu, Yibo, Huang, Wenhao, Sun, Huan, Su, Yu, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
by: Zou, Kai, et al.
Published: (2025)
by: Zou, Kai, et al.
Published: (2025)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
by: Zhang, Ge, et al.
Published: (2024)
by: Zhang, Ge, et al.
Published: (2024)
Ternary Gamma Semirings: From Neural Implementation to Categorical Foundations
by: Sun, Ruoqi
Published: (2026)
by: Sun, Ruoqi
Published: (2026)
Global coordination and trade‐off of grassland species traits and climatic drivers
by: Kuo Sun, et al.
Published: (2025)
by: Kuo Sun, et al.
Published: (2025)
QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
Spatial spillovers in trade agreement memberships: Does institutional proximity matter?
by: Renliang Liu, et al.
Published: (2024)
by: Renliang Liu, et al.
Published: (2024)
FitText: Evolving Agent Tool Ecologies via Memetic Retrieval
by: Zheng, Kyle, et al.
Published: (2026)
by: Zheng, Kyle, et al.
Published: (2026)
TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks
by: Jiang, Dongfu, et al.
Published: (2023)
by: Jiang, Dongfu, et al.
Published: (2023)
General-Reasoner: Advancing LLM Reasoning Across All Domains
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology
by: Sun, Yuxuan, et al.
Published: (2024)
by: Sun, Yuxuan, et al.
Published: (2024)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2025)
by: Ruan, Chi, et al.
Published: (2025)
CAPE: A Chinese Dataset for Appraisal-based Emotional Generation using Large Language Models
by: Liu, June M., et al.
Published: (2024)
by: Liu, June M., et al.
Published: (2024)
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
by: Hu, Kairui, et al.
Published: (2025)
by: Hu, Kairui, et al.
Published: (2025)
Programmable Sponge for Hydro‐Active Morphing Module with Light Weight and High‐Volume Change
by: Hajun Lee, et al.
Published: (2025)
by: Hajun Lee, et al.
Published: (2025)
Soil respiration at different time scales from 2000 to 2018 in forest ecosystems across China
by: Sun, Hongru, et al.
Published: (2021)
by: Sun, Hongru, et al.
Published: (2021)
Soil respiration at different time scales from 2000 to 2018 in forest ecosystems across China
by: Sun, Hongru, et al.
Published: (2022)
by: Sun, Hongru, et al.
Published: (2022)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
by: Zheng, Boyuan, et al.
Published: (2024)
by: Zheng, Boyuan, et al.
Published: (2024)
Stochastic Low-rank Tensor Bandits for Multi-dimensional Online Decision Making
by: Zhou, Jie, et al.
Published: (2020)
by: Zhou, Jie, et al.
Published: (2020)
Target-Guided Bayesian Flow Networks for Quantitatively Constrained CAD Generation
by: Zheng, Wenhao, et al.
Published: (2025)
by: Zheng, Wenhao, et al.
Published: (2025)
Fostering Natural Conversation in Large Language Models with NICO: a Natural Interactive COnversation dataset
by: Sun, Renliang, et al.
Published: (2024)
by: Sun, Renliang, et al.
Published: (2024)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
by: Ku, Max, et al.
Published: (2023)
by: Ku, Max, et al.
Published: (2023)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2026)
by: Ruan, Chi, et al.
Published: (2026)
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing
by: Zhang, Kai, et al.
Published: (2023)
by: Zhang, Kai, et al.
Published: (2023)
Age of Information Aided Intelligent Grant-Free Massive Access for Heterogeneous mMTC Traffic
by: Sun, Zhongwen, et al.
Published: (2025)
by: Sun, Zhongwen, et al.
Published: (2025)
A Panoramic Review on Intercalation‐Based Electrochemical Lithium Extraction From Salt Lakes: Mechanisms, Challenges, and Optimization Strategies
by: Shumin Wu, et al.
Published: (2026)
by: Shumin Wu, et al.
Published: (2026)
Research progress on quantum neural networks and quantum machine learning
by: Sun, Yifan, et al.
Published: (2026)
by: Sun, Yifan, et al.
Published: (2026)
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
by: Sun, Lingchen, et al.
Published: (2026)
by: Sun, Lingchen, et al.
Published: (2026)
Emergence is Overrated: AGI as an Archipelago of Experts
by: Kilov, Daniel
Published: (2026)
by: Kilov, Daniel
Published: (2026)
Wildfire and house prices: A synthetic control case study of Altadena (Jan 2025)
by: Sun, Yibo
Published: (2025)
by: Sun, Yibo
Published: (2025)
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
by: Wang, Yequan, et al.
Published: (2025)
by: Wang, Yequan, et al.
Published: (2025)
MANTIS: Interleaved Multi-Image Instruction Tuning
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents
by: Mo, Lingbo, et al.
Published: (2024)
by: Mo, Lingbo, et al.
Published: (2024)
The Arrival of AGI? When Expert Personas Exceed Expert Benchmarks
by: Mullens, Drake, et al.
Published: (2026)
by: Mullens, Drake, et al.
Published: (2026)
A bound-preserving Runge--Kutta discontinuous Galerkin method with compact stencils for hyperbolic conservation laws
by: Liu, Chen, et al.
Published: (2024)
by: Liu, Chen, et al.
Published: (2024)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
I-SHEEP: Self-Alignment of LLM from Scratch through an Iterative Self-Enhancement Paradigm
by: Liang, Yiming, et al.
Published: (2024)
by: Liang, Yiming, et al.
Published: (2024)
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
by: Li, Zhuofeng, et al.
Published: (2026)
by: Li, Zhuofeng, et al.
Published: (2026)
Similar Items
-
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024) -
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
by: Zou, Kai, et al.
Published: (2025) -
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024) -
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
by: Zhang, Ge, et al.
Published: (2024) -
Ternary Gamma Semirings: From Neural Implementation to Categorical Foundations
by: Sun, Ruoqi
Published: (2026)