Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Jiang, Yuxuan, Ferraro, Francis |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
por: Jiang, Yuxuan, et al.
Publicado: (2026)
por: Jiang, Yuxuan, et al.
Publicado: (2026)
DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
por: Jiang, Yuxuan, et al.
Publicado: (2025)
por: Jiang, Yuxuan, et al.
Publicado: (2025)
SAGA: A Participant-specific Examination of Story Alternatives and Goal Applicability for a Deeper Understanding of Complex Events
por: Vallurupalli, Sai, et al.
Publicado: (2024)
por: Vallurupalli, Sai, et al.
Publicado: (2024)
Explore the Reasoning Capability of LLMs in the Chess Testbed
por: Wang, Shu, et al.
Publicado: (2024)
por: Wang, Shu, et al.
Publicado: (2024)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
por: Kim, Jisu, et al.
Publicado: (2025)
por: Kim, Jisu, et al.
Publicado: (2025)
CoRE: Condition-based Reasoning for Identifying Outcome Variance in Complex Events
por: Vallurupalli, Sai, et al.
Publicado: (2025)
por: Vallurupalli, Sai, et al.
Publicado: (2025)
Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
por: Gabay, Adi, et al.
Publicado: (2026)
por: Gabay, Adi, et al.
Publicado: (2026)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
por: Sun, Jiaxing, et al.
Publicado: (2024)
por: Sun, Jiaxing, et al.
Publicado: (2024)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
por: Ruan, Zhiwen, et al.
Publicado: (2025)
por: Ruan, Zhiwen, et al.
Publicado: (2025)
World Models for Math Story Problems
por: Opedal, Andreas, et al.
Publicado: (2023)
por: Opedal, Andreas, et al.
Publicado: (2023)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
por: Dekoninck, Jasper, et al.
Publicado: (2026)
por: Dekoninck, Jasper, et al.
Publicado: (2026)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
por: Zhao, Yilun, et al.
Publicado: (2023)
por: Zhao, Yilun, et al.
Publicado: (2023)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
por: Hans, Abhimanyu, et al.
Publicado: (2024)
por: Hans, Abhimanyu, et al.
Publicado: (2024)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
por: O'Brien, Dayyán, et al.
Publicado: (2025)
por: O'Brien, Dayyán, et al.
Publicado: (2025)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
por: Guan, Xinyu, et al.
Publicado: (2025)
por: Guan, Xinyu, et al.
Publicado: (2025)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
por: Shelat, Shlok, et al.
Publicado: (2026)
por: Shelat, Shlok, et al.
Publicado: (2026)
Scheherazade: Evaluating Chain-of-Thought Math Reasoning in LLMs with Chain-of-Problems
por: Miner, Stephen, et al.
Publicado: (2024)
por: Miner, Stephen, et al.
Publicado: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
Right Is Not Enough: The Pitfalls of Outcome Supervision in Training LLMs for Math Reasoning
por: Guo, Jiaxing, et al.
Publicado: (2025)
por: Guo, Jiaxing, et al.
Publicado: (2025)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
por: Uddin, Md Nayem, et al.
Publicado: (2024)
por: Uddin, Md Nayem, et al.
Publicado: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
por: Li, Aochong Oliver, et al.
Publicado: (2025)
por: Li, Aochong Oliver, et al.
Publicado: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
por: Xu, Liang, et al.
Publicado: (2024)
por: Xu, Liang, et al.
Publicado: (2024)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
por: Lou, Siyu, et al.
Publicado: (2024)
por: Lou, Siyu, et al.
Publicado: (2024)
CoinMath: Harnessing the Power of Coding Instruction for Math LLMs
por: Wei, Chengwei, et al.
Publicado: (2024)
por: Wei, Chengwei, et al.
Publicado: (2024)
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs
por: Mukherjee, Sagnik, et al.
Publicado: (2025)
por: Mukherjee, Sagnik, et al.
Publicado: (2025)
TAG-EQA: Text-And-Graph for Event Question Answering via Structured Prompting Strategies
por: Kadam, Maithili, et al.
Publicado: (2025)
por: Kadam, Maithili, et al.
Publicado: (2025)
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains
por: Zhao, Yilun, et al.
Publicado: (2023)
por: Zhao, Yilun, et al.
Publicado: (2023)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
por: Liu, Hongwei, et al.
Publicado: (2024)
por: Liu, Hongwei, et al.
Publicado: (2024)
Reason to Rote: Rethinking Memorization in Reasoning
por: Du, Yupei, et al.
Publicado: (2025)
por: Du, Yupei, et al.
Publicado: (2025)
The CompMath-MCQ Dataset: Are LLMs Ready for Higher-Level Math?
por: Raimondi, Bianca, et al.
Publicado: (2026)
por: Raimondi, Bianca, et al.
Publicado: (2026)
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
por: Kassem, Aly M., et al.
Publicado: (2024)
por: Kassem, Aly M., et al.
Publicado: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
por: Roy, Tiasa Singha, et al.
Publicado: (2025)
por: Roy, Tiasa Singha, et al.
Publicado: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
por: Zeng, Liang, et al.
Publicado: (2024)
por: Zeng, Liang, et al.
Publicado: (2024)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
por: Balunović, Mislav, et al.
Publicado: (2025)
por: Balunović, Mislav, et al.
Publicado: (2025)
Mitigating Memorization in LLMs using Activation Steering
por: Suri, Manan, et al.
Publicado: (2025)
por: Suri, Manan, et al.
Publicado: (2025)
Beyond Memorization: The Challenge of Random Memory Access in Language Models
por: Zhu, Tongyao, et al.
Publicado: (2024)
por: Zhu, Tongyao, et al.
Publicado: (2024)
On Memorization of Large Language Models in Logical Reasoning
por: Xie, Chulin, et al.
Publicado: (2024)
por: Xie, Chulin, et al.
Publicado: (2024)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
por: Chen, Zui, et al.
Publicado: (2024)
por: Chen, Zui, et al.
Publicado: (2024)
If We May De-Presuppose: Robustly Verifying Claims through Presupposition-Free Question Decomposition
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
por: Dipta, Shubhashis Roy, et al.
Publicado: (2025)
Inductive Bias Extraction and Matching for LLM Prompts
por: Angel, Christian M., et al.
Publicado: (2025)
por: Angel, Christian M., et al.
Publicado: (2025)
Ejemplares similares
-
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
por: Jiang, Yuxuan, et al.
Publicado: (2026) -
DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models
por: Jiang, Yuxuan, et al.
Publicado: (2025) -
SAGA: A Participant-specific Examination of Story Alternatives and Goal Applicability for a Deeper Understanding of Complex Events
por: Vallurupalli, Sai, et al.
Publicado: (2024) -
Explore the Reasoning Capability of LLMs in the Chess Testbed
por: Wang, Shu, et al.
Publicado: (2024) -
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
por: Kim, Jisu, et al.
Publicado: (2025)