GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qintong, Cui, Leyang, Zhao, Xueliang, Kong, Lingpeng, Bi, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
by: Li, Qintong, et al.
Published: (2023)
by: Li, Qintong, et al.
Published: (2023)
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
by: Zhao, Xueliang, et al.
Published: (2024)
by: Zhao, Xueliang, et al.
Published: (2024)
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration
by: Li, Qintong, et al.
Published: (2024)
by: Li, Qintong, et al.
Published: (2024)
Scaling Reasoning without Attention
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
by: Zhao, Xueliang, et al.
Published: (2025)
by: Zhao, Xueliang, et al.
Published: (2025)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
MAGE: Machine-generated Text Detection in the Wild
by: Li, Yafu, et al.
Published: (2023)
by: Li, Yafu, et al.
Published: (2023)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
by: Jiang, Botian, et al.
Published: (2024)
by: Jiang, Botian, et al.
Published: (2024)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
by: Rahman, A M Muntasir, et al.
Published: (2024)
by: Rahman, A M Muntasir, et al.
Published: (2024)
Haste Makes Waste: Evaluating Planning Abilities of LLMs for Efficient and Feasible Multitasking with Time Constraints Between Actions
by: Wu, Zirui, et al.
Published: (2025)
by: Wu, Zirui, et al.
Published: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
Alleviating Hallucinations of Large Language Models through Induced Hallucinations
by: Zhang, Yue, et al.
Published: (2023)
by: Zhang, Yue, et al.
Published: (2023)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
by: Davoodi, Arash Gholami, et al.
Published: (2024)
by: Davoodi, Arash Gholami, et al.
Published: (2024)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
by: Xu, Zhiqiu, et al.
Published: (2026)
by: Xu, Zhiqiu, et al.
Published: (2026)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
by: Zhao, Xueliang, et al.
Published: (2024)
by: Zhao, Xueliang, et al.
Published: (2024)
How Well Do LLMs Handle Cantonese? Benchmarking Cantonese Capabilities of Large Language Models
by: Jiang, Jiyue, et al.
Published: (2024)
by: Jiang, Jiyue, et al.
Published: (2024)
MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation
by: Zeng, Zhongshen, et al.
Published: (2023)
by: Zeng, Zhongshen, et al.
Published: (2023)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
by: Zheng, Xiang, et al.
Published: (2026)
by: Zheng, Xiang, et al.
Published: (2026)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
Knowledge Verification to Nip Hallucination in the Bud
by: Wan, Fanqi, et al.
Published: (2024)
by: Wan, Fanqi, et al.
Published: (2024)
Jailbreaking as a Reward Misspecification Problem
by: Xie, Zhihui, et al.
Published: (2024)
by: Xie, Zhihui, et al.
Published: (2024)
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Spotting AI's Touch: Identifying LLM-Paraphrased Spans in Text
by: Li, Yafu, et al.
Published: (2024)
by: Li, Yafu, et al.
Published: (2024)
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations
by: Singh, Jyotika, et al.
Published: (2026)
by: Singh, Jyotika, et al.
Published: (2026)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
by: Li, Xin, et al.
Published: (2025)
by: Li, Xin, et al.
Published: (2025)
Retrieval is Accurate Generation
by: Cao, Bowen, et al.
Published: (2024)
by: Cao, Bowen, et al.
Published: (2024)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
Why Does the Effective Context Length of LLMs Fall Short?
by: An, Chenxin, et al.
Published: (2024)
by: An, Chenxin, et al.
Published: (2024)
U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs
by: Chernyshev, Konstantin, et al.
Published: (2024)
by: Chernyshev, Konstantin, et al.
Published: (2024)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening
by: Zhang, Ming, et al.
Published: (2026)
by: Zhang, Ming, et al.
Published: (2026)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
by: Majumdar, Ayan, et al.
Published: (2025)
by: Majumdar, Ayan, et al.
Published: (2025)
Dream-Coder 7B: An Open Diffusion Language Model for Code
by: Xie, Zhihui, et al.
Published: (2025)
by: Xie, Zhihui, et al.
Published: (2025)
Reasoning Does Not Necessarily Improve Role-Playing Ability
by: Feng, Xiachong, et al.
Published: (2025)
by: Feng, Xiachong, et al.
Published: (2025)
Linguistic Frameworks Go Toe-to-Toe at Neuro-Symbolic Language Modeling
by: Prange, Jakob, et al.
Published: (2021)
by: Prange, Jakob, et al.
Published: (2021)
Similar Items
-
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks
by: Li, Qintong, et al.
Published: (2023) -
DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
by: Zhao, Xueliang, et al.
Published: (2025) -
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
by: Zhao, Xueliang, et al.
Published: (2025) -
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models
by: Zhao, Xueliang, et al.
Published: (2024) -
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration
by: Li, Qintong, et al.
Published: (2024)