Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Gulati, Aryan, Miranda, Brando, Chen, Eric, Xia, Emily, Fronsdal, Kai, Dumont, Bruno, Obbad, Elyas, Koyejo, Sanmi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EvoPruneDeepTL: An Evolutionary Pruning Model for Transfer Learning based Deep Neural Networks
by: Poyatos, Javier, et al.
Published: (2022)
by: Poyatos, Javier, et al.
Published: (2022)
Alternate Loss Functions for Classification and Robust Regression Can Improve the Accuracy of Artificial Neural Networks
by: Noel, Mathew Mithra, et al.
Published: (2023)
by: Noel, Mathew Mithra, et al.
Published: (2023)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
Deterministic Event-Graph Substrates as World Models for Counterfactual Reasoning
by: Rovai, Fabio
Published: (2026)
by: Rovai, Fabio
Published: (2026)
Implementing Online Reinforcement Learning with Clustering Neural Networks
by: Smith, James E.
Published: (2024)
by: Smith, James E.
Published: (2024)
AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models
by: Perez, Miguel Angel Peñaloza, et al.
Published: (2025)
by: Perez, Miguel Angel Peñaloza, et al.
Published: (2025)
AiGAS-dEVL-RC: An Adaptive Growing Neural Gas Model for Recurrently Drifting Unsupervised Data Streams
by: Arostegi, Maria, et al.
Published: (2025)
by: Arostegi, Maria, et al.
Published: (2025)
Phase-Coded Memory and Morphological Resonance: A Next-Generation Retrieval-Augmented Generator Architecture
by: Saklakov, Denis V.
Published: (2025)
by: Saklakov, Denis V.
Published: (2025)
Incremental Bootstrapping and Classification of Structured Scenes in a Fuzzy Ontology
by: Buoncompagni, Luca, et al.
Published: (2024)
by: Buoncompagni, Luca, et al.
Published: (2024)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
by: Huo, Dongjie, et al.
Published: (2026)
by: Huo, Dongjie, et al.
Published: (2026)
How much do LLMs learn from negative examples?
by: Hamdan, Shadi, et al.
Published: (2025)
by: Hamdan, Shadi, et al.
Published: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
by: Qi, Dekang, et al.
Published: (2026)
by: Qi, Dekang, et al.
Published: (2026)
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
by: Liu, Bowen, et al.
Published: (2026)
by: Liu, Bowen, et al.
Published: (2026)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)
by: Berman, Shmuel, et al.
Published: (2024)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
Loss shaping enhances exact gradient learning with Eventprop in spiking neural networks
by: Nowotny, Thomas, et al.
Published: (2022)
by: Nowotny, Thomas, et al.
Published: (2022)
Optimizing Genetic Algorithms Using the Binomial Distribution
by: Cicirello, Vincent A.
Published: (2024)
by: Cicirello, Vincent A.
Published: (2024)
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
by: Sharma, Aditya, et al.
Published: (2025)
by: Sharma, Aditya, et al.
Published: (2025)
Heterogeneous LLM Methods for Ontology Learning (Few-Shot Prompting, Ensemble Typing, and Attention-Based Taxonomies)
by: Beliaeva, Aleksandra, et al.
Published: (2025)
by: Beliaeva, Aleksandra, et al.
Published: (2025)
Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge
by: Kricheli, Joshua Shay, et al.
Published: (2024)
by: Kricheli, Joshua Shay, et al.
Published: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
Modeling Local Search Metaheuristics Using Markov Decision Processes
by: Ruiz-Torrubiano, Rubén
Published: (2024)
by: Ruiz-Torrubiano, Rubén
Published: (2024)
Physics-Driven AI Correction in Laser Absorption Sensing Quantification
by: Kang, Ruiyuan, et al.
Published: (2024)
by: Kang, Ruiyuan, et al.
Published: (2024)
Reservoir Computing with Evolved Critical Neural Cellular Automata
by: Pontes-Filho, Sidney, et al.
Published: (2025)
by: Pontes-Filho, Sidney, et al.
Published: (2025)
On measuring grounding and generalizing grounding problems
by: Quigley, Daniel, et al.
Published: (2025)
by: Quigley, Daniel, et al.
Published: (2025)
An effective Genetic Programming Hyper-Heuristic for Uncertain Agile Satellite Scheduling
by: Chen, Yuning, et al.
Published: (2026)
by: Chen, Yuning, et al.
Published: (2026)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
LTL Verification of Memoryful Neural Agents
by: Hosseini, Mehran, et al.
Published: (2025)
by: Hosseini, Mehran, et al.
Published: (2025)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
by: Koh, Hyunseo, et al.
Published: (2026)
by: Koh, Hyunseo, et al.
Published: (2026)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
by: Haque, Md. Asraful, et al.
Published: (2026)
by: Haque, Md. Asraful, et al.
Published: (2026)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
by: Ma, Yueen, et al.
Published: (2024)
by: Ma, Yueen, et al.
Published: (2024)
A Landmark-Aware Visual Navigation Dataset
by: Johnson, Faith, et al.
Published: (2024)
by: Johnson, Faith, et al.
Published: (2024)
Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
by: Burns, Thomas F, et al.
Published: (2024)
by: Burns, Thomas F, et al.
Published: (2024)
Approximating Discrimination Within Models When Faced With Several Non-Binary Sensitive Attributes
by: Bian, Yijun, et al.
Published: (2024)
by: Bian, Yijun, et al.
Published: (2024)
Does Machine Bring in Extra Bias in Learning? Approximating Fairness in Models Promptly
by: Bian, Yijun, et al.
Published: (2024)
by: Bian, Yijun, et al.
Published: (2024)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
by: Menon, Anjali R., et al.
Published: (2025)
by: Menon, Anjali R., et al.
Published: (2025)
Similar Items
-
EvoPruneDeepTL: An Evolutionary Pruning Model for Transfer Learning based Deep Neural Networks
by: Poyatos, Javier, et al.
Published: (2022) -
Alternate Loss Functions for Classification and Robust Regression Can Improve the Accuracy of Artificial Neural Networks
by: Noel, Mathew Mithra, et al.
Published: (2023) -
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024) -
Deterministic Event-Graph Substrates as World Models for Counterfactual Reasoning
by: Rovai, Fabio
Published: (2026) -
Implementing Online Reinforcement Learning with Clustering Neural Networks
by: Smith, James E.
Published: (2024)