MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yubo, Ma, Xueguang, Zhang, Ge, Ni, Yuansheng, Chandra, Abhranil, Guo, Shiguang, Ren, Weiming, Arulraj, Aaran, He, Xuan, Jiang, Ziyan, Li, Tianle, Ku, Max, Wang, Kai, Zhuang, Alex, Fan, Rongqi, Yue, Xiang, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
by: He, Xuan, et al.
Published: (2024)
by: He, Xuan, et al.
Published: (2024)
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023)
by: Wang, Yubo, et al.
Published: (2023)
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
ImagenHub: Standardizing the evaluation of conditional image generation models
by: Ku, Max, et al.
Published: (2023)
by: Ku, Max, et al.
Published: (2023)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
by: Chen, Jiacheng, et al.
Published: (2024)
by: Chen, Jiacheng, et al.
Published: (2024)
Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
by: Chandra, Abhranil, et al.
Published: (2025)
by: Chandra, Abhranil, et al.
Published: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
by: KJ, Sankalp, et al.
Published: (2025)
by: KJ, Sankalp, et al.
Published: (2025)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
by: Ma, Wentao, et al.
Published: (2025)
by: Ma, Wentao, et al.
Published: (2025)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
VISA: Retrieval Augmented Generation with Visual Source Attribution
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
New methods to compute the generalized chi-square distribution
by: Das, Abhranil
Published: (2024)
by: Das, Abhranil
Published: (2024)
DisProtEdit: Exploring Disentangled Representations for Multi-Attribute Protein Editing
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
General-Reasoner: Advancing LLM Reasoning Across All Domains
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Towards a Periodic Table of Computer System Design Principles
by: Arulraj, Joy
Published: (2025)
by: Arulraj, Joy
Published: (2025)
Are We Done with MMLU?
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
MANTIS: Interleaved Multi-Image Instruction Tuning
by: Jiang, Dongfu, et al.
Published: (2024)
by: Jiang, Dongfu, et al.
Published: (2024)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
by: Ku, Max, et al.
Published: (2023)
by: Ku, Max, et al.
Published: (2023)
Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2025)
by: Ruan, Chi, et al.
Published: (2025)
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
Long-context LLMs Struggle with Long In-context Learning
by: Li, Tianle, et al.
Published: (2024)
by: Li, Tianle, et al.
Published: (2024)
Methods to integrate multinormals and compute classification measures
by: Das, Abhranil, et al.
Published: (2020)
by: Das, Abhranil, et al.
Published: (2020)
Unifying Multimodal Retrieval via Document Screenshot Embedding
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
Shopping MMLU: A Massive Multi-Task Online Shopping Benchmark for Large Language Models
by: Jin, Yilun, et al.
Published: (2024)
by: Jin, Yilun, et al.
Published: (2024)
MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed
by: Zhan, Jiaqi Samantha, et al.
Published: (2025)
by: Zhan, Jiaqi Samantha, et al.
Published: (2025)
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
StructLM: Towards Building Generalist Models for Structured Knowledge Grounding
by: Zhuang, Alex, et al.
Published: (2024)
by: Zhuang, Alex, et al.
Published: (2024)
Mobile-MMLU: A Mobile Intelligence Language Understanding Benchmark
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
by: Bsharat, Sondos Mahmoud, et al.
Published: (2025)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
LLM Collaboration With Multi-Agent Reinforcement Learning
by: Liu, Shuo, et al.
Published: (2025)
by: Liu, Shuo, et al.
Published: (2025)
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
by: Narsupalli, Yaswanth, et al.
Published: (2024)
by: Narsupalli, Yaswanth, et al.
Published: (2024)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
by: Cai, Songcheng, et al.
Published: (2026)
by: Cai, Songcheng, et al.
Published: (2026)
Similar Items
-
VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation
by: He, Xuan, et al.
Published: (2024) -
GenAI Arena: An Open Evaluation Platform for Generative Models
by: Jiang, Dongfu, et al.
Published: (2024) -
Augmenting Black-box LLMs with Medical Textbooks for Biomedical Question Answering
by: Wang, Yubo, et al.
Published: (2023) -
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
by: Jiang, Ziyan, et al.
Published: (2024) -
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
by: Yue, Xiang, et al.
Published: (2024)