MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Zhongshen, Liu, Yinhong, Wan, Yingjia, Li, Jingyao, Chen, Pengguang, Dai, Jianbo, Yao, Yuxuan, Xu, Rongwu, Qi, Zehan, Zhao, Wanru, Shen, Linling, Lu, Jianqiao, Tan, Haochen, Chen, Yukang, Zhang, Hao, Shi, Zhan, Wang, Bailin, Guo, Zhijiang, Jia, Jiaya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
by: Zeng, Zhongshen, et al.
Published: (2023)
by: Zeng, Zhongshen, et al.
Published: (2023)
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
by: Wan, Yingjia, et al.
Published: (2025)
by: Wan, Yingjia, et al.
Published: (2025)
VLPose: Bridging the Domain Gap in Pose Estimation with Language-Vision Tuning
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation
by: Li, Jingyao, et al.
Published: (2023)
by: Li, Jingyao, et al.
Published: (2023)
MOODv2: Masked Image Modeling for Out-of-Distribution Detection
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
by: Li, Jingyao, et al.
Published: (2023)
by: Li, Jingyao, et al.
Published: (2023)
AutoPSV: Automated Process-Supervised Verifier
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
RoboCoder: Robotic Learning from Basic Skills to General Tasks with Large Language Models
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
Knowledge Conflicts for LLMs: A Survey
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
DebateQA: Evaluating Question Answering on Debatable Knowledge
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
FormalAlign: Automated Alignment Evaluation for Autoformalization
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
Long$^2$RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall
by: Qi, Zehan, et al.
Published: (2024)
by: Qi, Zehan, et al.
Published: (2024)
GridMask Data Augmentation
by: Chen, Pengguang, et al.
Published: (2020)
by: Chen, Pengguang, et al.
Published: (2020)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
by: Yang, Senqiao, et al.
Published: (2024)
by: Yang, Senqiao, et al.
Published: (2024)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
by: Lai, Xin, et al.
Published: (2024)
by: Lai, Xin, et al.
Published: (2024)
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation
by: Dai, Jianbo, et al.
Published: (2024)
by: Dai, Jianbo, et al.
Published: (2024)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
Meta-Auxiliary Learning for Micro-Expression Recognition
by: Wang, Jingyao, et al.
Published: (2024)
by: Wang, Jingyao, et al.
Published: (2024)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Exploring Chinese Humor Generation: A Study on Two-Part Allegorical Sayings
by: Xu, Rongwu
Published: (2024)
by: Xu, Rongwu
Published: (2024)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
Aligning with Logic: Measuring, Evaluating and Improving Logical Preference Consistency in Large Language Models
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
AwesomeMeta+: A Mixed-Prototyping Meta-Learning System Supporting AI Application Design Anywhere
by: Wang, Jingyao, et al.
Published: (2023)
by: Wang, Jingyao, et al.
Published: (2023)
Process-Driven Autoformalization in Lean 4
by: Lu, Jianqiao, et al.
Published: (2024)
by: Lu, Jianqiao, et al.
Published: (2024)
LISA: Reasoning Segmentation via Large Language Model
by: Lai, Xin, et al.
Published: (2023)
by: Lai, Xin, et al.
Published: (2023)
OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
by: Chen, Yukang, et al.
Published: (2023)
by: Chen, Yukang, et al.
Published: (2023)
When Efficiency Backfires: Cascading LLMs Trigger Cascade Failure under Adversarial Attack
by: Sun, Zehan, et al.
Published: (2026)
by: Sun, Zehan, et al.
Published: (2026)
Who Is a Better Bargainer?
by: Linling Geng
Published: (2024)
by: Linling Geng
Published: (2024)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
by: Chen, Zaoyu, et al.
Published: (2026)
by: Chen, Zaoyu, et al.
Published: (2026)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
QuickLLaMA: Query-aware Inference Acceleration for Large Language Models
by: Li, Jingyao, et al.
Published: (2024)
by: Li, Jingyao, et al.
Published: (2024)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
by: Liu, Shaoteng, et al.
Published: (2024)
by: Liu, Shaoteng, et al.
Published: (2024)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
by: Wang, Chengyao, et al.
Published: (2025)
by: Wang, Chengyao, et al.
Published: (2025)
Tempo: Confidentiality Preservation in Cloud-Based Neural Network Training
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
by: Wei, Shaohang, et al.
Published: (2025)
by: Wei, Shaohang, et al.
Published: (2025)
Fear of happiness and emotion regulation: Implications for personal and relational well‐being
by: Huixian Acacia Lee, et al.
Published: (2025)
by: Huixian Acacia Lee, et al.
Published: (2025)
Improving Preference Extraction In LLMs By Identifying Latent Knowledge Through Classifying Probes
by: Maiya, Sharan, et al.
Published: (2025)
by: Maiya, Sharan, et al.
Published: (2025)
Similar Items
-
MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
by: Zeng, Zhongshen, et al.
Published: (2023) -
FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs
by: Wan, Yingjia, et al.
Published: (2025) -
VLPose: Bridging the Domain Gap in Pose Estimation with Language-Vision Tuning
by: Li, Jingyao, et al.
Published: (2024) -
TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation
by: Li, Jingyao, et al.
Published: (2023) -
MOODv2: Masked Image Modeling for Out-of-Distribution Detection
by: Li, Jingyao, et al.
Published: (2024)