GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Mirzadeh, Iman, Alizadeh, Keivan, Shahrokhi, Hooman, Tuzel, Oncel, Bengio, Samy, Farajtabar, Mehrdad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025)
by: Shojaee, Parshin, et al.
Published: (2025)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
by: Joudaki, Amir, et al.
Published: (2025)
by: Joudaki, Amir, et al.
Published: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023)
by: Alizadeh, Keivan, et al.
Published: (2023)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024)
by: Samragh, Mohammad, et al.
Published: (2024)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
by: Alizadeh, Keivan, et al.
Published: (2024)
by: Alizadeh, Keivan, et al.
Published: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
by: Alizadeh, Keivan, et al.
Published: (2026)
by: Alizadeh, Keivan, et al.
Published: (2026)
Computational Bottlenecks of Training Small-scale Large Language Models
by: Ashkboos, Saleh, et al.
Published: (2024)
by: Ashkboos, Saleh, et al.
Published: (2024)
RL for Reasoning by Adaptively Revealing Rationales
by: Amani, Mohammad Hossein, et al.
Published: (2025)
by: Amani, Mohammad Hossein, et al.
Published: (2025)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2025)
by: Li, Jeffrey, et al.
Published: (2025)
Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
by: Chegini, Atoosa, et al.
Published: (2025)
by: Chegini, Atoosa, et al.
Published: (2025)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
by: Chegini, Atoosa, et al.
Published: (2024)
by: Chegini, Atoosa, et al.
Published: (2024)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
by: Zhu, Hanlin, et al.
Published: (2025)
by: Zhu, Hanlin, et al.
Published: (2025)
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
by: Abbe, Emmanuel, et al.
Published: (2024)
by: Abbe, Emmanuel, et al.
Published: (2024)
Conformal Prediction Sets for Deep Generative Models via Reduction to Conformal Regression
by: Shahrokhi, Hooman, et al.
Published: (2025)
by: Shahrokhi, Hooman, et al.
Published: (2025)
Pretraining with hierarchical memories: separating long-tail and common knowledge
by: Pouransari, Hadi, et al.
Published: (2025)
by: Pouransari, Hadi, et al.
Published: (2025)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
by: Vemulapalli, Raviteja, et al.
Published: (2023)
by: Vemulapalli, Raviteja, et al.
Published: (2023)
Jailbreaking Large Language Models with Symbolic Mathematics
by: Bethany, Emet, et al.
Published: (2024)
by: Bethany, Emet, et al.
Published: (2024)
Conformal Thinking: Risk Control for Reasoning on a Compute Budget
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
by: Zan, Lei, et al.
Published: (2025)
by: Zan, Lei, et al.
Published: (2025)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
by: Yu, Zhouliang, et al.
Published: (2025)
by: Yu, Zhouliang, et al.
Published: (2025)
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
by: Shrestha, Safal, et al.
Published: (2026)
by: Shrestha, Safal, et al.
Published: (2026)
Neuro-Symbolic Artificial Intelligence: Towards Improving the Reasoning Abilities of Large Language Models
by: Yang, Xiao-Wen, et al.
Published: (2025)
by: Yang, Xiao-Wen, et al.
Published: (2025)
Advanced Weakly-Supervised Formula Exploration for Neuro-Symbolic Mathematical Reasoning
by: Wu, Yuxuan, et al.
Published: (2025)
by: Wu, Yuxuan, et al.
Published: (2025)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
by: Shao, Zhihong, et al.
Published: (2024)
by: Shao, Zhihong, et al.
Published: (2024)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
by: Wang, Haoxiang, et al.
Published: (2023)
by: Wang, Haoxiang, et al.
Published: (2023)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
Learning to Reason for Hallucination Span Detection
by: Su, Hsuan, et al.
Published: (2025)
by: Su, Hsuan, et al.
Published: (2025)
ULLER: A Unified Language for Learning and Reasoning
by: van Krieken, Emile, et al.
Published: (2024)
by: van Krieken, Emile, et al.
Published: (2024)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
by: Zhao, Zilong, et al.
Published: (2024)
by: Zhao, Zilong, et al.
Published: (2024)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
by: Jiang, Weisen, et al.
Published: (2023)
by: Jiang, Weisen, et al.
Published: (2023)
Efficient Causal Graph Discovery Using Large Language Models
by: Jiralerspong, Thomas, et al.
Published: (2024)
by: Jiralerspong, Thomas, et al.
Published: (2024)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
by: Zhang, Zheng
Published: (2025)
by: Zhang, Zheng
Published: (2025)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
by: Cinquin, Tristan, et al.
Published: (2025)
by: Cinquin, Tristan, et al.
Published: (2025)
On the Limits of Self-Improving in Large Language Models: The Singularity Is Not Near Without Symbolic Model Synthesis
by: Zenil, Hector
Published: (2026)
by: Zenil, Hector
Published: (2026)
Towards a Mechanistic Understanding of Propositional Logical Reasoning in Large Language Models
by: Chen, Danchun, et al.
Published: (2026)
by: Chen, Danchun, et al.
Published: (2026)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
by: Yao, Jian, et al.
Published: (2025)
by: Yao, Jian, et al.
Published: (2025)
Similar Items
-
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
by: Shojaee, Parshin, et al.
Published: (2025) -
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
by: Joudaki, Amir, et al.
Published: (2025) -
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023) -
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
by: Samragh, Mohammad, et al.
Published: (2024) -
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
by: Alizadeh, Keivan, et al.
Published: (2024)