A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rahman, Roussel, Mishra, Aashwin Ananda |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reversing the Lens: Using Explainable AI to Understand Human Expertise
von: Rahman, Roussel, et al.
Veröffentlicht: (2025)
von: Rahman, Roussel, et al.
Veröffentlicht: (2025)
A Small Math Model: Recasting Strategy Choice Theory in an LLM-Inspired Architecture
von: Rahman, Roussel, et al.
Veröffentlicht: (2025)
von: Rahman, Roussel, et al.
Veröffentlicht: (2025)
Large Language Models in Numberland: A Quick Test of Their Numerical Reasoning Abilities
von: Rahman, Roussel
Veröffentlicht: (2025)
von: Rahman, Roussel
Veröffentlicht: (2025)
Uncertainty Quantification via Stable Distribution Propagation
von: Petersen, Felix, et al.
Veröffentlicht: (2024)
von: Petersen, Felix, et al.
Veröffentlicht: (2024)
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)
Reinforcing Numerical Reasoning in LLMs for Tabular Prediction via Structural Priors
von: Cai, Pengxiang, et al.
Veröffentlicht: (2025)
von: Cai, Pengxiang, et al.
Veröffentlicht: (2025)
Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs
von: Vijayaraj, Raj Krishnan
Veröffentlicht: (2025)
von: Vijayaraj, Raj Krishnan
Veröffentlicht: (2025)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025)
von: Cinquin, Tristan, et al.
Veröffentlicht: (2025)
EDGE: A Theoretical Framework for Misconception-Aware Adaptive Learning
von: Verma, Ananda Prakash
Veröffentlicht: (2025)
von: Verma, Ananda Prakash
Veröffentlicht: (2025)
When Can LLMs Learn to Reason with Weak Supervision?
von: Rahman, Salman, et al.
Veröffentlicht: (2026)
von: Rahman, Salman, et al.
Veröffentlicht: (2026)
On The Fragility of Benchmark Contamination Detection in Reasoning Models
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
von: Zhang, Zheng
Veröffentlicht: (2025)
von: Zhang, Zheng
Veröffentlicht: (2025)
Language Models Do Not Embed Numbers Continuously
von: Davies, Alex O., et al.
Veröffentlicht: (2025)
von: Davies, Alex O., et al.
Veröffentlicht: (2025)
A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization
von: Liu, Haoxin, et al.
Veröffentlicht: (2024)
von: Liu, Haoxin, et al.
Veröffentlicht: (2024)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
von: Feng, Guhao, et al.
Veröffentlicht: (2024)
von: Feng, Guhao, et al.
Veröffentlicht: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
Preventing Curriculum Collapse in Self-Evolving Reasoning Systems
von: Mishra, Vaibhav
Veröffentlicht: (2026)
von: Mishra, Vaibhav
Veröffentlicht: (2026)
Towards single-shot coherent imaging via overlap-free ptychography
von: Hoidn, Oliver, et al.
Veröffentlicht: (2026)
von: Hoidn, Oliver, et al.
Veröffentlicht: (2026)
Finding the Cracks: Improving LLMs Reasoning with Paraphrastic Probing and Consistency Verification
von: Shi, Weili, et al.
Veröffentlicht: (2026)
von: Shi, Weili, et al.
Veröffentlicht: (2026)
Probing Knowledge Holes in Unlearned LLMs
von: Ko, Myeongseob, et al.
Veröffentlicht: (2025)
von: Ko, Myeongseob, et al.
Veröffentlicht: (2025)
Evidence for Limited Metacognition in LLMs
von: Ackerman, Christopher
Veröffentlicht: (2025)
von: Ackerman, Christopher
Veröffentlicht: (2025)
seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
von: Ramezanali, Mohammad, et al.
Veröffentlicht: (2025)
QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs
von: Mishra, Himanshu, et al.
Veröffentlicht: (2026)
von: Mishra, Himanshu, et al.
Veröffentlicht: (2026)
Exam Readiness Index (ERI): A Theoretical Framework for a Composite, Explainable Index
von: Verma, Ananda Prakash
Veröffentlicht: (2025)
von: Verma, Ananda Prakash
Veröffentlicht: (2025)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
von: McGovern, Hope, et al.
Veröffentlicht: (2026)
Eliciting Numerical Predictive Distributions of LLMs Without Autoregression
von: Piskorz, Julianna, et al.
Veröffentlicht: (2026)
von: Piskorz, Julianna, et al.
Veröffentlicht: (2026)
Red-teaming Activation Probes using Prompted LLMs
von: Blandfort, Phil, et al.
Veröffentlicht: (2025)
von: Blandfort, Phil, et al.
Veröffentlicht: (2025)
Domain Knowledge Guided Bayesian Optimization For Autonomous Alignment Of Complex Scientific Instruments
von: Mishra, Aashwin, et al.
Veröffentlicht: (2026)
von: Mishra, Aashwin, et al.
Veröffentlicht: (2026)
Fast and Accurate Probing of In-Training LLMs' Downstream Performances
von: Liu, Zhichen, et al.
Veröffentlicht: (2026)
von: Liu, Zhichen, et al.
Veröffentlicht: (2026)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
any4: Learned 4-bit Numeric Representation for LLMs
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2025)
LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2024)
von: Rahman, Md Hafizur, et al.
Veröffentlicht: (2024)
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
von: Korbak, Tomek, et al.
Veröffentlicht: (2025)
Probing the Trajectories of Reasoning Traces in Large Language Models
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
von: Ballon, Marthe, et al.
Veröffentlicht: (2026)
FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
von: Li, Zhuoyun, et al.
Veröffentlicht: (2026)
von: Li, Zhuoyun, et al.
Veröffentlicht: (2026)
Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents
von: K, Deeraj S, et al.
Veröffentlicht: (2026)
von: K, Deeraj S, et al.
Veröffentlicht: (2026)
FLEX: Feature Importance from Layered Counterfactual Explanations
von: Keshtmand, Nawid, et al.
Veröffentlicht: (2025)
von: Keshtmand, Nawid, et al.
Veröffentlicht: (2025)
Learning to Solve Resource-Constrained Project Scheduling Problems with Duration Uncertainty using Graph Neural Networks
von: Infantes, Guillaume, et al.
Veröffentlicht: (2025)
von: Infantes, Guillaume, et al.
Veröffentlicht: (2025)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
von: Fernandez, Nigel, et al.
Veröffentlicht: (2025)
von: Fernandez, Nigel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reversing the Lens: Using Explainable AI to Understand Human Expertise
von: Rahman, Roussel, et al.
Veröffentlicht: (2025) -
A Small Math Model: Recasting Strategy Choice Theory in an LLM-Inspired Architecture
von: Rahman, Roussel, et al.
Veröffentlicht: (2025) -
Large Language Models in Numberland: A Quick Test of Their Numerical Reasoning Abilities
von: Rahman, Roussel
Veröffentlicht: (2025) -
Uncertainty Quantification via Stable Distribution Propagation
von: Petersen, Felix, et al.
Veröffentlicht: (2024) -
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs
von: Zhou, Yifan, et al.
Veröffentlicht: (2025)