The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
Fuente:
arXiv
Saved in:
| Main Authors: | Bertolazzi, Leonardo, Mondorf, Philipp, Plank, Barbara, Bernardi, Raffaella |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
by: Bertolazzi, Leonardo, et al.
Published: (2025)
by: Bertolazzi, Leonardo, et al.
Published: (2025)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
by: Rabern, Brian, et al.
Published: (2026)
by: Rabern, Brian, et al.
Published: (2026)
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
by: Bertolazzi, Leonardo, et al.
Published: (2024)
by: Bertolazzi, Leonardo, et al.
Published: (2024)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Tracing Uncertainty in Language Model "Reasoning"
by: Grünefeld, Nils, et al.
Published: (2026)
by: Grünefeld, Nils, et al.
Published: (2026)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
by: Orth, Jasmin, et al.
Published: (2025)
by: Orth, Jasmin, et al.
Published: (2025)
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
by: Cheng, Xinyuan, et al.
Published: (2026)
by: Cheng, Xinyuan, et al.
Published: (2026)
Teaching Small Language Models to Learn Logic through Meta-Learning
by: Bertolazzi, Leonardo, et al.
Published: (2025)
by: Bertolazzi, Leonardo, et al.
Published: (2025)
Arithmetic with Language Models: from Memorization to Computation
by: Maltoni, Davide, et al.
Published: (2023)
by: Maltoni, Davide, et al.
Published: (2023)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
by: Mondorf, Philipp, et al.
Published: (2025)
by: Mondorf, Philipp, et al.
Published: (2025)
Language Models Fail to Introspect About Their Knowledge of Language
by: Song, Siyuan, et al.
Published: (2025)
by: Song, Siyuan, et al.
Published: (2025)
Human-Calibrated Automated Testing and Validation of Generative Language Models
by: Sudjianto, Agus, et al.
Published: (2024)
by: Sudjianto, Agus, et al.
Published: (2024)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026)
by: Lee, Seongyun, et al.
Published: (2026)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Probing for Arithmetic Errors in Language Models
by: Sun, Yucheng, et al.
Published: (2025)
by: Sun, Yucheng, et al.
Published: (2025)
Measuring what Matters: Construct Validity in Large Language Model Benchmarks
by: Bean, Andrew M., et al.
Published: (2025)
by: Bean, Andrew M., et al.
Published: (2025)
Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models
by: Pope, Quintin, et al.
Published: (2026)
by: Pope, Quintin, et al.
Published: (2026)
How LLMs Fail to Support Fact-Checking
by: Proma, Adiba Mahbub, et al.
Published: (2025)
by: Proma, Adiba Mahbub, et al.
Published: (2025)
Self-training Language Models for Arithmetic Reasoning
by: Kadlčík, Marek, et al.
Published: (2024)
by: Kadlčík, Marek, et al.
Published: (2024)
When Valid Signals Fail: Regime Boundaries Between LLM Features and RL Trading Policies
by: Yang, Zhengzhe
Published: (2026)
by: Yang, Zhengzhe
Published: (2026)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
by: Lin, Shiyin
Published: (2025)
by: Lin, Shiyin
Published: (2025)
Prompting or Fine-tuning? Exploring Large Language Models for Causal Graph Validation
by: Susanti, Yuni, et al.
Published: (2024)
by: Susanti, Yuni, et al.
Published: (2024)
Selecting Language Models for Social Science: Start Small, Start Open, and Validate
by: Stoltz, Dustin S., et al.
Published: (2026)
by: Stoltz, Dustin S., et al.
Published: (2026)
Mechanistic Behavior Editing of Language Models
by: Singh, Joykirat, et al.
Published: (2024)
by: Singh, Joykirat, et al.
Published: (2024)
Larger Language Models Don't Care How You Think: Why Chain-of-Thought Prompting Fails in Subjective Tasks
by: Chochlakis, Georgios, et al.
Published: (2024)
by: Chochlakis, Georgios, et al.
Published: (2024)
Latent Trajectory Dynamics in Large Language Models: A Manifold Evolution Framework with Empirical Validation
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
Integrating Large Language Models and Knowledge Graphs for Extraction and Validation of Textual Test Data
by: De Santis, Antonio, et al.
Published: (2024)
by: De Santis, Antonio, et al.
Published: (2024)
Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports
by: Niu, Chuang, et al.
Published: (2024)
by: Niu, Chuang, et al.
Published: (2024)
Bridging Language Gaps: Enhancing Few-Shot Language Adaptation
by: Borchert, Philipp, et al.
Published: (2025)
by: Borchert, Philipp, et al.
Published: (2025)
Mechanistic Indicators of Understanding in Large Language Models
by: Beckmann, Pierre, et al.
Published: (2025)
by: Beckmann, Pierre, et al.
Published: (2025)
Mechanistic Origin of Moral Indifference in Language Models
by: Li, Lingyu, et al.
Published: (2026)
by: Li, Lingyu, et al.
Published: (2026)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
by: Xu, Shanshan, et al.
Published: (2025)
by: Xu, Shanshan, et al.
Published: (2025)
Modular Arithmetic: Language Models Solve Math Digit by Digit
by: Baeumel, Tanja, et al.
Published: (2025)
by: Baeumel, Tanja, et al.
Published: (2025)
Understanding When Tree of Thoughts Succeeds: Larger Models Excel in Generation, Not Discrimination
by: Chen, Qiqi, et al.
Published: (2024)
by: Chen, Qiqi, et al.
Published: (2024)
From Voices to Validity: Leveraging Large Language Models (LLMs) for Textual Analysis of Policy Stakeholder Interviews
by: Liu, Alex, et al.
Published: (2023)
by: Liu, Alex, et al.
Published: (2023)
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
by: Zhao, Boxiang, et al.
Published: (2026)
by: Zhao, Boxiang, et al.
Published: (2026)
Similar Items
-
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024) -
How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
by: Bertolazzi, Leonardo, et al.
Published: (2025) -
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
by: Mondorf, Philipp, et al.
Published: (2024) -
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
by: Rabern, Brian, et al.
Published: (2026) -
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences
by: Bertolazzi, Leonardo, et al.
Published: (2024)