I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Camposampiero, Giacomo, Hersche, Michael, Wattenhofer, Roger, Sebastian, Abu, Rahimi, Abbas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning
von: Hersche, Michael, et al.
Veröffentlicht: (2024)
von: Hersche, Michael, et al.
Veröffentlicht: (2024)
Towards Learning Abductive Reasoning using VSA Distributed Representations
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2024)
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2024)
Scalable Evaluation and Neural Models for Compositional Generalization
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2024)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2024)
Limits of Transformer Language Models on Learning to Compose Algorithms
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
Probabilistic Abduction for Visual Abstract Reasoning via Learning Rules in Vector-symbolic Architectures
von: Hersche, Michael, et al.
Veröffentlicht: (2024)
von: Hersche, Michael, et al.
Veröffentlicht: (2024)
Terminating Differentiable Tree Experts
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
On the Role of Noise in Factorizers for Disentangling Distributed Representations
von: Karunaratne, Geethan, et al.
Veröffentlicht: (2024)
von: Karunaratne, Geethan, et al.
Veröffentlicht: (2024)
A foundation model with multi-variate parallel attention to generate neuronal activity
von: Carzaniga, Francesco, et al.
Veröffentlicht: (2025)
von: Carzaniga, Francesco, et al.
Veröffentlicht: (2025)
Locally Coherent Parallel Decoding in Diffusion Language Models
von: Hersche, Michael, et al.
Veröffentlicht: (2026)
von: Hersche, Michael, et al.
Veröffentlicht: (2026)
Soft-Masked Diffusion Language Models
von: Hersche, Michael, et al.
Veröffentlicht: (2025)
von: Hersche, Michael, et al.
Veröffentlicht: (2025)
A-I-RAVEN and I-RAVEN-Mesh: Two New Benchmarks for Abstract Visual Reasoning
von: Małkiński, Mikołaj, et al.
Veröffentlicht: (2024)
von: Małkiński, Mikołaj, et al.
Veröffentlicht: (2024)
A Theoretical Analysis of Test-Driven Code Generation
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2025)
PUZZLES: A Benchmark for Neural Algorithmic Reasoning
von: Estermann, Benjamin, et al.
Veröffentlicht: (2024)
von: Estermann, Benjamin, et al.
Veröffentlicht: (2024)
Evaluating the Robustness of Analogical Reasoning in Large Language Models
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
von: Lewis, Martha, et al.
Veröffentlicht: (2024)
Thompson Sampling via Fine-Tuning of LLMs
von: Menet, Nicolas, et al.
Veröffentlicht: (2025)
von: Menet, Nicolas, et al.
Veröffentlicht: (2025)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
FLIP Reasoning Challenge
von: Plesner, Andreas, et al.
Veröffentlicht: (2025)
von: Plesner, Andreas, et al.
Veröffentlicht: (2025)
Kernel Approximation using Analog In-Memory Computing
von: Büchel, Julian, et al.
Veröffentlicht: (2024)
von: Büchel, Julian, et al.
Veröffentlicht: (2024)
Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2026)
Benchmarking Positional Encodings for GNNs and Graph Transformers
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
von: Grötschla, Florian, et al.
Veröffentlicht: (2024)
The Case for Cleaner Biosignals: High-fidelity Neural Compressor Enables Transfer from Cleaner iEEG to Noisier EEG
von: Carzaniga, Francesco Stefano, et al.
Veröffentlicht: (2025)
von: Carzaniga, Francesco Stefano, et al.
Veröffentlicht: (2025)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
von: Mirzadeh, Iman, et al.
Veröffentlicht: (2024)
von: Mirzadeh, Iman, et al.
Veröffentlicht: (2024)
CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
von: Zan, Lei, et al.
Veröffentlicht: (2025)
von: Zan, Lei, et al.
Veröffentlicht: (2025)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
von: Pawar, Pranav, et al.
Veröffentlicht: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
Reasoning Effort and Problem Complexity: A Scaling Analysis in LLMs
von: Estermann, Benjamin, et al.
Veröffentlicht: (2025)
von: Estermann, Benjamin, et al.
Veröffentlicht: (2025)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
von: Zhao, Zilong, et al.
Veröffentlicht: (2024)
von: Zhao, Zilong, et al.
Veröffentlicht: (2024)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Beyond Interpolation: Extrapolative Reasoning with Reinforcement Learning and Graph Neural Networks
von: Grillo, Niccolò, et al.
Veröffentlicht: (2025)
von: Grillo, Niccolò, et al.
Veröffentlicht: (2025)
Benchmarking Reasoning Robustness in Large Language Models
von: Yu, Tong, et al.
Veröffentlicht: (2025)
von: Yu, Tong, et al.
Veröffentlicht: (2025)
GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models
von: Wu, Zixuan, et al.
Veröffentlicht: (2025)
von: Wu, Zixuan, et al.
Veröffentlicht: (2025)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Parameterized Argumentation-based Reasoning Tasks for Benchmarking Generative Language Models
von: Steging, Cor, et al.
Veröffentlicht: (2025)
von: Steging, Cor, et al.
Veröffentlicht: (2025)
On the Expressive Power of GNNs for Boolean Satisfiability
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
von: Peltonen, Saku, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025) -
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning
von: Hersche, Michael, et al.
Veröffentlicht: (2024) -
Towards Learning Abductive Reasoning using VSA Distributed Representations
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2024) -
Scalable Evaluation and Neural Models for Compositional Generalization
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025) -
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages
von: Terzić, Aleksandar, et al.
Veröffentlicht: (2024)