MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Pengfei, Zhang, Fanrui, Peng, Xiaopeng, Xu, Zhaopan, Ai, Jiaxin, Qiu, Yansheng, Li, Chuanhao, Li, Zhen, Li, Ming, Feng, Yukang, Sun, Jianwen, Zhang, Haoquan, Li, Zizhen, Mao, Xiaofeng, Zhao, Wangbo, Wang, Kai, Chang, Xiaojun, Shao, Wenqi, You, Yang, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
Closing the Expression Gap in LLM Instructions via Socratic Questioning
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences
by: Li, Zizhen, et al.
Published: (2026)
by: Li, Zizhen, et al.
Published: (2026)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)
by: Feng, Yukang, et al.
Published: (2025)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
by: Li, Zizhen, et al.
Published: (2025)
by: Li, Zizhen, et al.
Published: (2025)
World Craft: Agentic Framework to Create Visualizable Worlds via Text
by: Sun, Jianwen, et al.
Published: (2026)
by: Sun, Jianwen, et al.
Published: (2026)
ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
AI Idea Bench 2025: AI Research Idea Generation Benchmark
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
AutoBG: A Board Game Design Assistant with Interactive Ideation, Iterative Rulebook Generation, and Individualized Feedback
by: Li, Zizhen, et al.
Published: (2026)
by: Li, Zizhen, et al.
Published: (2026)
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
by: Zhang, Fanrui, et al.
Published: (2025)
by: Zhang, Fanrui, et al.
Published: (2025)
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
by: Qiu, Yansheng, et al.
Published: (2025)
by: Qiu, Yansheng, et al.
Published: (2025)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
Sekai: A Video Dataset towards World Exploration
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
IA-T2I: Internet-Augmented Text-to-Image Generation
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
by: Zhou, Pengfei, et al.
Published: (2024)
by: Zhou, Pengfei, et al.
Published: (2024)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
LaGen: Towards Autoregressive LiDAR Scene Generation
by: Zhou, Sizhuo, et al.
Published: (2025)
by: Zhou, Sizhuo, et al.
Published: (2025)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Yume-1.5: A Text-Controlled Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
Yume: An Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
Einstellungstest MDK
by: Easy Quizzz
Published: (2026)
by: Easy Quizzz
Published: (2026)
PyVision: Agentic Vision with Dynamic Tooling
by: Zhao, Shitian, et al.
Published: (2025)
by: Zhao, Shitian, et al.
Published: (2025)
PyVision-RL: Forging Open Agentic Vision Models via RL
by: Zhao, Shitian, et al.
Published: (2026)
by: Zhao, Shitian, et al.
Published: (2026)
Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
WorldMark: A Unified Benchmark Suite for Interactive Video World Models
by: Xu, Xiaojie, et al.
Published: (2026)
by: Xu, Xiaojie, et al.
Published: (2026)
Deciphering the Acetabular Labrum's Cellular Atlas: MDK Inhibition as a Novel Therapeutic Method for Developmental Dysplasia of the Hip
by: Runze Yang, et al.
Published: (2025)
by: Runze Yang, et al.
Published: (2025)
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
by: Li, Zhen, et al.
Published: (2026)
by: Li, Zhen, et al.
Published: (2026)
SVBench: Evaluation of Video Generation Models on Social Reasoning
by: Peng, Wenshuo, et al.
Published: (2025)
by: Peng, Wenshuo, et al.
Published: (2025)
Edge-aware Decoding for Neural Asymmetric Routing
by: Liang, Li, et al.
Published: (2026)
by: Liang, Li, et al.
Published: (2026)
REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
by: Wang, Ziqiao, et al.
Published: (2025)
by: Wang, Ziqiao, et al.
Published: (2025)
Efficient High-Resolution Visual Representation Learning with State Space Model for Human Pose Estimation
by: Zhang, Hao, et al.
Published: (2024)
by: Zhang, Hao, et al.
Published: (2024)
Similar Items
-
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
by: Zhou, Pengfei, et al.
Published: (2025) -
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025) -
Closing the Expression Gap in LLM Instructions via Socratic Questioning
by: Sun, Jianwen, et al.
Published: (2025) -
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
by: Sun, Jianwen, et al.
Published: (2025) -
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)