MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Taghanaki, Saeid Asgari, Khani, Aliasgahr, Khasahmadi, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
by: Taghanaki, Saeid Asgari, et al.
Published: (2025)
by: Taghanaki, Saeid Asgari, et al.
Published: (2025)
TExplain: Explaining Learned Visual Features via Pre-trained (Frozen) Language Models
by: Taghanaki, Saeid Asgari, et al.
Published: (2023)
by: Taghanaki, Saeid Asgari, et al.
Published: (2023)
SLiMe: Segment Like Me
by: Khani, Aliasghar, et al.
Published: (2023)
by: Khani, Aliasghar, et al.
Published: (2023)
Detecting Generative Parroting through Overfitting Masked Autoencoders
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
by: Yuan, Yu, et al.
Published: (2024)
by: Yuan, Yu, et al.
Published: (2024)
Reasoning in Transformers -- Mitigating Spurious Correlations and Reasoning Shortcuts
by: Enström, Daniel, et al.
Published: (2024)
by: Enström, Daniel, et al.
Published: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
by: Rakhsha, Amin, et al.
Published: (2025)
by: Rakhsha, Amin, et al.
Published: (2025)
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
by: Wang, Wentian, et al.
Published: (2024)
by: Wang, Wentian, et al.
Published: (2024)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
by: Zhou, Yuqing, et al.
Published: (2024)
by: Zhou, Yuqing, et al.
Published: (2024)
Short-circuiting Shortcuts: Mechanistic Investigation of Shortcuts in Text Classification
by: Eshuijs, Leon, et al.
Published: (2025)
by: Eshuijs, Leon, et al.
Published: (2025)
Disentangled PET Lesion Segmentation
by: Gatsak, Tanya, et al.
Published: (2024)
by: Gatsak, Tanya, et al.
Published: (2024)
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
by: Yan, Lecheng, et al.
Published: (2026)
by: Yan, Lecheng, et al.
Published: (2026)
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
by: Cekinmez, Jasin, et al.
Published: (2025)
by: Cekinmez, Jasin, et al.
Published: (2025)
Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing
by: Sheikhi, Saeid
Published: (2026)
by: Sheikhi, Saeid
Published: (2026)
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding
by: Cheshmi, Seyyed Saeid, et al.
Published: (2026)
by: Cheshmi, Seyyed Saeid, et al.
Published: (2026)
Deep Semantic Segmentation of Natural and Medical Images: A Review
by: Taghanaki, Saeid Asgari, et al.
Published: (2019)
by: Taghanaki, Saeid Asgari, et al.
Published: (2019)
Learning to Reason in LLMs by Expectation Maximization
by: Lee, Junghyun, et al.
Published: (2025)
by: Lee, Junghyun, et al.
Published: (2025)
Evaluating LLMs' Reasoning Over Ordered Procedural Steps
by: Anika, Adrita, et al.
Published: (2025)
by: Anika, Adrita, et al.
Published: (2025)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
Higher-Order Knowledge Representations for Agentic Scientific Reasoning
by: Stewart, Isabella A., et al.
Published: (2026)
by: Stewart, Isabella A., et al.
Published: (2026)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
by: Bihani, Geetanjali, et al.
Published: (2024)
by: Bihani, Geetanjali, et al.
Published: (2024)
Seeing to Generalize: How Visual Data Corrects Binding Shortcuts
by: Buzeta, Nicolas, et al.
Published: (2026)
by: Buzeta, Nicolas, et al.
Published: (2026)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
by: Dang, Quy-Anh, et al.
Published: (2025)
by: Dang, Quy-Anh, et al.
Published: (2025)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
by: Contestabile, Matilde, et al.
Published: (2025)
by: Contestabile, Matilde, et al.
Published: (2025)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
The Reliability Paradox: Exploring How Shortcut Learning Undermines Language Model Calibration
by: Bihani, Geetanjali, et al.
Published: (2024)
by: Bihani, Geetanjali, et al.
Published: (2024)
Where Norms and References Collide: Evaluating LLMs on Normative Reasoning
by: Abrams, Mitchell, et al.
Published: (2026)
by: Abrams, Mitchell, et al.
Published: (2026)
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
by: Xuan, Weihao, et al.
Published: (2025)
by: Xuan, Weihao, et al.
Published: (2025)
IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge
by: Abdelaal, Ali, et al.
Published: (2026)
by: Abdelaal, Ali, et al.
Published: (2026)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
by: Ni, Ruikang, et al.
Published: (2024)
by: Ni, Ruikang, et al.
Published: (2024)
Failure Modes of LLMs for Causal Reasoning on Narratives
by: Yamin, Khurram, et al.
Published: (2024)
by: Yamin, Khurram, et al.
Published: (2024)
Demystifying Long Chain-of-Thought Reasoning in LLMs
by: Yeo, Edward, et al.
Published: (2025)
by: Yeo, Edward, et al.
Published: (2025)
Probabilistic Reasoning with LLMs for k-anonymity Estimation
by: Zheng, Jonathan, et al.
Published: (2025)
by: Zheng, Jonathan, et al.
Published: (2025)
Prompt Engineering a Prompt Engineer
by: Ye, Qinyuan, et al.
Published: (2023)
by: Ye, Qinyuan, et al.
Published: (2023)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
by: Yue, Murong, et al.
Published: (2024)
by: Yue, Murong, et al.
Published: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
How to Determine the Preferred Image Distribution of a Black-Box Vision-Language Model?
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
by: Taghanaki, Saeid Asgari, et al.
Published: (2024)
Learning to Correct for QA Reasoning with Black-box LLMs
by: Kim, Jaehyung, et al.
Published: (2024)
by: Kim, Jaehyung, et al.
Published: (2024)
Similar Items
-
Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
by: Taghanaki, Saeid Asgari, et al.
Published: (2025) -
TExplain: Explaining Learned Visual Features via Pre-trained (Frozen) Language Models
by: Taghanaki, Saeid Asgari, et al.
Published: (2023) -
SLiMe: Segment Like Me
by: Khani, Aliasghar, et al.
Published: (2023) -
Detecting Generative Parroting through Overfitting Masked Autoencoders
by: Taghanaki, Saeid Asgari, et al.
Published: (2024) -
Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models
by: Yuan, Yu, et al.
Published: (2024)