Beyond Memorization: Distinguishing between Reductive and Epistemic Reasoning in LLMs using Classic Logic Puzzles
Fuente:
arXiv
Guardado en:
| Autores principales: | Gabay, Adi, Stanovsky, Gabriel, Peterfreund, Liat |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
por: Simhi, Adi, et al.
Publicado: (2025)
por: Simhi, Adi, et al.
Publicado: (2025)
Handling SQL Nulls with Two-Valued Logic
por: Libkin, Leonid, et al.
Publicado: (2020)
por: Libkin, Leonid, et al.
Publicado: (2020)
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
por: Jiang, Yuxuan, et al.
Publicado: (2024)
por: Jiang, Yuxuan, et al.
Publicado: (2024)
On Memorization of Large Language Models in Logical Reasoning
por: Xie, Chulin, et al.
Publicado: (2024)
por: Xie, Chulin, et al.
Publicado: (2024)
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
por: Lior, Gili, et al.
Publicado: (2025)
por: Lior, Gili, et al.
Publicado: (2025)
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition
por: Goldstein, Ariel, et al.
Publicado: (2024)
por: Goldstein, Ariel, et al.
Publicado: (2024)
The State and Fate of Summarization Datasets: A Survey
por: Dahan, Noam, et al.
Publicado: (2024)
por: Dahan, Noam, et al.
Publicado: (2024)
Comparing Humans and Models on a Similar Scale: Towards Cognitive Gender Bias Evaluation in Coreference Resolution
por: Lior, Gili, et al.
Publicado: (2023)
por: Lior, Gili, et al.
Publicado: (2023)
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
por: Kim, Jisu, et al.
Publicado: (2025)
por: Kim, Jisu, et al.
Publicado: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
por: Itzhak, Itay, et al.
Publicado: (2025)
por: Itzhak, Itay, et al.
Publicado: (2025)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
por: Park, Jungsoo, et al.
Publicado: (2025)
por: Park, Jungsoo, et al.
Publicado: (2025)
Looking Beyond The Top-1: Transformers Determine Top Tokens In Order
por: Lioubashevski, Daria, et al.
Publicado: (2024)
por: Lioubashevski, Daria, et al.
Publicado: (2024)
Unveiling Over-Memorization in Finetuning LLMs for Reasoning Tasks
por: Ruan, Zhiwen, et al.
Publicado: (2025)
por: Ruan, Zhiwen, et al.
Publicado: (2025)
Beyond Benchmarks: On The False Promise of AI Regulation
por: Stanovsky, Gabriel, et al.
Publicado: (2025)
por: Stanovsky, Gabriel, et al.
Publicado: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks
por: Salido, Eva Sánchez, et al.
Publicado: (2025)
por: Salido, Eva Sánchez, et al.
Publicado: (2025)
Mitigating Memorization in LLMs using Activation Steering
por: Suri, Manan, et al.
Publicado: (2025)
por: Suri, Manan, et al.
Publicado: (2025)
Leveraging Collection-Wide Similarities for Unsupervised Document Structure Extraction
por: Lior, Gili, et al.
Publicado: (2024)
por: Lior, Gili, et al.
Publicado: (2024)
Leveraging Digitized Newspapers to Collect Summarization Data in Low-Resource Languages
por: Dahan, Noam, et al.
Publicado: (2025)
por: Dahan, Noam, et al.
Publicado: (2025)
In-Context Learning on a Budget: A Case Study in Token Classification
por: Berger, Uri, et al.
Publicado: (2024)
por: Berger, Uri, et al.
Publicado: (2024)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
por: Jahara, Fatima, et al.
Publicado: (2025)
por: Jahara, Fatima, et al.
Publicado: (2025)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
por: Itzhak, Itay, et al.
Publicado: (2026)
por: Itzhak, Itay, et al.
Publicado: (2026)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
por: Sun, Jiaxing, et al.
Publicado: (2024)
por: Sun, Jiaxing, et al.
Publicado: (2024)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
por: Badola, Kartikeya, et al.
Publicado: (2025)
por: Badola, Kartikeya, et al.
Publicado: (2025)
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
por: Chen, Jiangjie, et al.
Publicado: (2025)
por: Chen, Jiangjie, et al.
Publicado: (2025)
Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles
por: Shah, Kulin, et al.
Publicado: (2024)
por: Shah, Kulin, et al.
Publicado: (2024)
AutoLogi: Automated Generation of Logic Puzzles for Evaluating Reasoning Abilities of Large Language Models
por: Zhu, Qin, et al.
Publicado: (2025)
por: Zhu, Qin, et al.
Publicado: (2025)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
por: Shelat, Shlok, et al.
Publicado: (2026)
por: Shelat, Shlok, et al.
Publicado: (2026)
RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like Reasoning
por: Chan, Jason, et al.
Publicado: (2024)
por: Chan, Jason, et al.
Publicado: (2024)
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
por: Chen, Guizhen, et al.
Publicado: (2025)
por: Chen, Guizhen, et al.
Publicado: (2025)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
por: Uddin, Md Nayem, et al.
Publicado: (2024)
por: Uddin, Md Nayem, et al.
Publicado: (2024)
Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMs
por: Hans, Abhimanyu, et al.
Publicado: (2024)
por: Hans, Abhimanyu, et al.
Publicado: (2024)
Memorization vs. Reasoning: Updating LLMs with New Knowledge
por: Li, Aochong Oliver, et al.
Publicado: (2025)
por: Li, Aochong Oliver, et al.
Publicado: (2025)
On the Expressiveness of Languages for Querying Property Graphs in Relational Databases
por: Rotschield, Hadar, et al.
Publicado: (2025)
por: Rotschield, Hadar, et al.
Publicado: (2025)
Database Theory in Action: From Inexpressibility to Efficiency in GQL's Order-Constrained Paths
por: Rotschield, Hadar, et al.
Publicado: (2025)
por: Rotschield, Hadar, et al.
Publicado: (2025)
Towards Cross-Model Efficiency in SQL/PGQ
por: Rotschield, Hadar, et al.
Publicado: (2025)
por: Rotschield, Hadar, et al.
Publicado: (2025)
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
por: Zhang, Lan, et al.
Publicado: (2025)
por: Zhang, Lan, et al.
Publicado: (2025)
Quantifying In-Context Reasoning Effects and Memorization Effects in LLMs
por: Lou, Siyu, et al.
Publicado: (2024)
por: Lou, Siyu, et al.
Publicado: (2024)
PromptSuite: A Task-Agnostic Framework for Multi-Prompt Generation
por: Habba, Eliya, et al.
Publicado: (2025)
por: Habba, Eliya, et al.
Publicado: (2025)
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation
por: Iluz, Bar, et al.
Publicado: (2024)
por: Iluz, Bar, et al.
Publicado: (2024)
Ejemplares similares
-
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
por: Simhi, Adi, et al.
Publicado: (2025) -
Handling SQL Nulls with Two-Valued Logic
por: Libkin, Leonid, et al.
Publicado: (2020) -
Beyond Math: Stories as a Testbed for Memorization-Constrained Reasoning in LLMs
por: Jiang, Yuxuan, et al.
Publicado: (2024) -
On Memorization of Large Language Models in Logical Reasoning
por: Xie, Chulin, et al.
Publicado: (2024) -
Comparing the Framing Effect in Humans and LLMs on Naturally Occurring Texts
por: Lior, Gili, et al.
Publicado: (2025)