Saved in:
| Main Author: | Chojecki, Przemyslaw |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.13764 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Psychometric Tests for AI Agents and Their Moduli Space
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Model Science: getting serious about verification, explanation and control of AI systems
by: Biecek, Przemyslaw, et al.
Published: (2025)
by: Biecek, Przemyslaw, et al.
Published: (2025)
On the Mathematical Impossibility of Safe Universal Approximators
by: Yao, Jasper
Published: (2025)
by: Yao, Jasper
Published: (2025)
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
by: Fan, Jingxuan, et al.
Published: (2024)
by: Fan, Jingxuan, et al.
Published: (2024)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
by: Hao, Yuren, et al.
Published: (2025)
by: Hao, Yuren, et al.
Published: (2025)
Beyond Backpropagation: Exploring Innovative Algorithms for Energy-Efficient Deep Neural Network Training
by: Spyra, Przemysław
Published: (2025)
by: Spyra, Przemysław
Published: (2025)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
by: Yu, Zhouliang, et al.
Published: (2025)
by: Yu, Zhouliang, et al.
Published: (2025)
FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks
by: Thomas, Nishal, et al.
Published: (2026)
by: Thomas, Nishal, et al.
Published: (2026)
The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer
by: Chen, Tianhua
Published: (2026)
by: Chen, Tianhua
Published: (2026)
The Nature of Mathematical Modeling and Probabilistic Optimization Engineering in Generative AI
by: Li, Fulu
Published: (2024)
by: Li, Fulu
Published: (2024)
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
by: Spyrison, Nicholas, et al.
Published: (2022)
by: Spyrison, Nicholas, et al.
Published: (2022)
Attributions All the Way Down? The Metagame of Interpretability
by: Baniecki, Hubert, et al.
Published: (2026)
by: Baniecki, Hubert, et al.
Published: (2026)
Parity, Sensitivity, and Transformers
by: Kozachinskiy, Alexander, et al.
Published: (2026)
by: Kozachinskiy, Alexander, et al.
Published: (2026)
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
by: Roggeveen, James V., et al.
Published: (2025)
by: Roggeveen, James V., et al.
Published: (2025)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
by: Yao, Jian, et al.
Published: (2025)
by: Yao, Jian, et al.
Published: (2025)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
by: Singh, Kunal, et al.
Published: (2025)
by: Singh, Kunal, et al.
Published: (2025)
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
by: Buckley, Warren, et al.
Published: (2023)
by: Buckley, Warren, et al.
Published: (2023)
Position: Explain to Question not to Justify
by: Biecek, Przemyslaw, et al.
Published: (2024)
by: Biecek, Przemyslaw, et al.
Published: (2024)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
by: Camposampiero, Giacomo, et al.
Published: (2025)
by: Camposampiero, Giacomo, et al.
Published: (2025)
AIDE: AI-Driven Exploration in the Space of Code
by: Jiang, Zhengyao, et al.
Published: (2025)
by: Jiang, Zhengyao, et al.
Published: (2025)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
by: Zimmer, Max, et al.
Published: (2026)
by: Zimmer, Max, et al.
Published: (2026)
REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop
by: Rybak, Patryk, et al.
Published: (2026)
by: Rybak, Patryk, et al.
Published: (2026)
AI Agents as Universal Task Solvers
by: Achille, Alessandro, et al.
Published: (2025)
by: Achille, Alessandro, et al.
Published: (2025)
Universal AI maximizes Variational Empowerment
by: Hayashi, Yusuke, et al.
Published: (2025)
by: Hayashi, Yusuke, et al.
Published: (2025)
Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction
by: Wahl, Ari, et al.
Published: (2026)
by: Wahl, Ari, et al.
Published: (2026)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
by: Li, Chengpeng, et al.
Published: (2024)
by: Li, Chengpeng, et al.
Published: (2024)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
by: Lohn, Evan, et al.
Published: (2024)
by: Lohn, Evan, et al.
Published: (2024)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
BAID: A Benchmark for Bias Assessment of AI Detectors
by: Basu, Priyam, et al.
Published: (2025)
by: Basu, Priyam, et al.
Published: (2025)
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Democratizing AI scientists using ToolUniverse
by: Gao, Shanghua, et al.
Published: (2025)
by: Gao, Shanghua, et al.
Published: (2025)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
by: Polowczyk, Agnieszka, et al.
Published: (2025)
by: Polowczyk, Agnieszka, et al.
Published: (2025)
FreSh: Frequency Shifting for Accelerated Neural Representation Learning
by: Kania, Adam, et al.
Published: (2024)
by: Kania, Adam, et al.
Published: (2024)
Exploration of the Rashomon Set Assists Trustworthy Explanations for Medical Data
by: Kobylińska, Katarzyna, et al.
Published: (2023)
by: Kobylińska, Katarzyna, et al.
Published: (2023)
Make Interval Bound Propagation great again
by: Krukowski, Patryk, et al.
Published: (2024)
by: Krukowski, Patryk, et al.
Published: (2024)
Bounding Evidence and Estimating Log-Likelihood in VAE
by: Struski, Łukasz, et al.
Published: (2022)
by: Struski, Łukasz, et al.
Published: (2022)
Similar Items
-
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025) -
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025) -
Psychometric Tests for AI Agents and Their Moduli Space
by: Chojecki, Przemyslaw
Published: (2025) -
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
by: Chojecki, Przemyslaw
Published: (2025) -
Model Science: getting serious about verification, explanation and control of AI systems
by: Biecek, Przemyslaw, et al.
Published: (2025)