Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Leo, Zhang, Jonathan, Chua, Sean, Kim, Spencer, Zhu, Kevin, O'Brien, Sean, Sharma, Vasu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
by: Aluru, Aayush, et al.
Published: (2025)
by: Aluru, Aayush, et al.
Published: (2025)
WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
by: Agarwal, Mrinal, et al.
Published: (2025)
by: Agarwal, Mrinal, et al.
Published: (2025)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
by: Begin, James, et al.
Published: (2025)
by: Begin, James, et al.
Published: (2025)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
by: Gupta, Abhay, et al.
Published: (2025)
by: Gupta, Abhay, et al.
Published: (2025)
The Geometry of Harmfulness in LLMs through Subconcept Probing
by: Shah, McNair, et al.
Published: (2025)
by: Shah, McNair, et al.
Published: (2025)
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
by: Nandan, Advey, et al.
Published: (2025)
by: Nandan, Advey, et al.
Published: (2025)
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
by: Wen, Athena, et al.
Published: (2025)
by: Wen, Athena, et al.
Published: (2025)
CBMAS: Cognitive Behavioral Modeling via Activation Steering
by: Ismail, Ahmed H., et al.
Published: (2026)
by: Ismail, Ahmed H., et al.
Published: (2026)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
by: Baek, Shaun, et al.
Published: (2025)
by: Baek, Shaun, et al.
Published: (2025)
Improving LLM Abilities in Idiomatic Translation
by: Donthi, Sundesh, et al.
Published: (2024)
by: Donthi, Sundesh, et al.
Published: (2024)
Large Language Models and Mathematical Reasoning Failures
by: Boye, Johan, et al.
Published: (2025)
by: Boye, Johan, et al.
Published: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Mathematical Computation and Reasoning Errors by Large Language Models
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
Key-Point-Driven Mathematical Reasoning Distillation of Large Language Model
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
by: Kim, Hyeonwoo, et al.
Published: (2024)
by: Kim, Hyeonwoo, et al.
Published: (2024)
Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
A Survey on Mathematical Reasoning and Optimization with Large Language Models
by: Forootani, Ali
Published: (2025)
by: Forootani, Ali
Published: (2025)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
by: Wang, Zhijie
Published: (2026)
by: Wang, Zhijie
Published: (2026)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
by: Amjad, Husnain, et al.
Published: (2026)
by: Amjad, Husnain, et al.
Published: (2026)
On Reasoning Strength Planning in Large Reasoning Models
by: Sheng, Leheng, et al.
Published: (2025)
by: Sheng, Leheng, et al.
Published: (2025)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
A Survey on Large Language Models for Mathematical Reasoning
by: Wang, Peng-Yuan, et al.
Published: (2025)
by: Wang, Peng-Yuan, et al.
Published: (2025)
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
by: Rufail, Andrew, et al.
Published: (2025)
by: Rufail, Andrew, et al.
Published: (2025)
Enhancing Japanese Large Language Models with Reasoning Vectors
by: Oguchi, Carolina Minami, et al.
Published: (2025)
by: Oguchi, Carolina Minami, et al.
Published: (2025)
SwiftSolve: A Self-Iterative, Complexity-Aware Multi-Agent Framework for Competitive Programming
by: Singh, Adhyayan Veer, et al.
Published: (2025)
by: Singh, Adhyayan Veer, et al.
Published: (2025)
Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning
by: Zhao, Jun, et al.
Published: (2024)
by: Zhao, Jun, et al.
Published: (2024)
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
by: Chaturvedi, Isha, et al.
Published: (2025)
by: Chaturvedi, Isha, et al.
Published: (2025)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
by: Karaman, Batuhan K., et al.
Published: (2026)
by: Karaman, Batuhan K., et al.
Published: (2026)
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models
by: Wang, Junjian, et al.
Published: (2026)
by: Wang, Junjian, et al.
Published: (2026)
Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
by: Sun, Zhishen, et al.
Published: (2025)
by: Sun, Zhishen, et al.
Published: (2025)
A Beautiful Mind: Principles and Strategies for AI-Augmented Human Reasoning
by: Koon, Sean
Published: (2025)
by: Koon, Sean
Published: (2025)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
by: Li, Xiaoyuan, et al.
Published: (2024)
by: Li, Xiaoyuan, et al.
Published: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
by: Mirzadeh, Iman, et al.
Published: (2024)
by: Mirzadeh, Iman, et al.
Published: (2024)
CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
by: Zan, Lei, et al.
Published: (2025)
by: Zan, Lei, et al.
Published: (2025)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
Similar Items
-
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
by: Aluru, Aayush, et al.
Published: (2025) -
WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
by: Agarwal, Mrinal, et al.
Published: (2025) -
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
by: Xiong, Lang, et al.
Published: (2025) -
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025) -
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
by: Begin, James, et al.
Published: (2025)