Gespeichert in:
| 1. Verfasser: | Chojecki, Przemyslaw |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2512.13764 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Improving AI Agents through Self-Play
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
The Geometry of Benchmarks: A New Path Toward AGI
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
Psychometric Tests for AI Agents and Their Moduli Space
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
von: Chojecki, Przemyslaw
Veröffentlicht: (2025)
Model Science: getting serious about verification, explanation and control of AI systems
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2025)
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2025)
On the Mathematical Impossibility of Safe Universal Approximators
von: Yao, Jasper
Veröffentlicht: (2025)
von: Yao, Jasper
Veröffentlicht: (2025)
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
von: Fan, Jingxuan, et al.
Veröffentlicht: (2024)
von: Fan, Jingxuan, et al.
Veröffentlicht: (2024)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
Beyond Backpropagation: Exploring Innovative Algorithms for Energy-Efficient Deep Neural Network Training
von: Spyra, Przemysław
Veröffentlicht: (2025)
von: Spyra, Przemysław
Veröffentlicht: (2025)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks
von: Thomas, Nishal, et al.
Veröffentlicht: (2026)
von: Thomas, Nishal, et al.
Veröffentlicht: (2026)
The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer
von: Chen, Tianhua
Veröffentlicht: (2026)
von: Chen, Tianhua
Veröffentlicht: (2026)
The Nature of Mathematical Modeling and Probabilistic Optimization Engineering in Generative AI
von: Li, Fulu
Veröffentlicht: (2024)
von: Li, Fulu
Veröffentlicht: (2024)
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
von: Spyrison, Nicholas, et al.
Veröffentlicht: (2022)
von: Spyrison, Nicholas, et al.
Veröffentlicht: (2022)
Attributions All the Way Down? The Metagame of Interpretability
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
von: Baniecki, Hubert, et al.
Veröffentlicht: (2026)
Parity, Sensitivity, and Transformers
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2026)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2026)
HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class
von: Roggeveen, James V., et al.
Veröffentlicht: (2025)
von: Roggeveen, James V., et al.
Veröffentlicht: (2025)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
von: Yao, Jian, et al.
Veröffentlicht: (2025)
von: Yao, Jian, et al.
Veröffentlicht: (2025)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
von: Buckley, Warren, et al.
Veröffentlicht: (2023)
von: Buckley, Warren, et al.
Veröffentlicht: (2023)
Position: Explain to Question not to Justify
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2024)
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2024)
I-RAVEN-X: Benchmarking Generalization and Robustness of Analogical and Mathematical Reasoning in Large Language and Reasoning Models
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
von: Camposampiero, Giacomo, et al.
Veröffentlicht: (2025)
AIDE: AI-Driven Exploration in the Space of Code
von: Jiang, Zhengyao, et al.
Veröffentlicht: (2025)
von: Jiang, Zhengyao, et al.
Veröffentlicht: (2025)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
von: Zimmer, Max, et al.
Veröffentlicht: (2026)
von: Zimmer, Max, et al.
Veröffentlicht: (2026)
REBEL: Hidden Knowledge Recovery via Evolutionary-Based Evaluation Loop
von: Rybak, Patryk, et al.
Veröffentlicht: (2026)
von: Rybak, Patryk, et al.
Veröffentlicht: (2026)
AI Agents as Universal Task Solvers
von: Achille, Alessandro, et al.
Veröffentlicht: (2025)
von: Achille, Alessandro, et al.
Veröffentlicht: (2025)
Universal AI maximizes Variational Empowerment
von: Hayashi, Yusuke, et al.
Veröffentlicht: (2025)
von: Hayashi, Yusuke, et al.
Veröffentlicht: (2025)
Monocular 3D Object Position Estimation with VLMs for Human-Robot Interaction
von: Wahl, Ari, et al.
Veröffentlicht: (2026)
von: Wahl, Ari, et al.
Veröffentlicht: (2026)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
von: Lohn, Evan, et al.
Veröffentlicht: (2024)
von: Lohn, Evan, et al.
Veröffentlicht: (2024)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
von: Liu, Xinyu, et al.
Veröffentlicht: (2026)
von: Liu, Xinyu, et al.
Veröffentlicht: (2026)
BAID: A Benchmark for Bias Assessment of AI Detectors
von: Basu, Priyam, et al.
Veröffentlicht: (2025)
von: Basu, Priyam, et al.
Veröffentlicht: (2025)
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Solving a Research Problem in Mathematical Statistics with AI Assistance
von: Dobriban, Edgar
Veröffentlicht: (2025)
von: Dobriban, Edgar
Veröffentlicht: (2025)
Democratizing AI scientists using ToolUniverse
von: Gao, Shanghua, et al.
Veröffentlicht: (2025)
von: Gao, Shanghua, et al.
Veröffentlicht: (2025)
Memory Self-Regeneration: Uncovering Hidden Knowledge in Unlearned Models
von: Polowczyk, Agnieszka, et al.
Veröffentlicht: (2025)
von: Polowczyk, Agnieszka, et al.
Veröffentlicht: (2025)
FreSh: Frequency Shifting for Accelerated Neural Representation Learning
von: Kania, Adam, et al.
Veröffentlicht: (2024)
von: Kania, Adam, et al.
Veröffentlicht: (2024)
Exploration of the Rashomon Set Assists Trustworthy Explanations for Medical Data
von: Kobylińska, Katarzyna, et al.
Veröffentlicht: (2023)
von: Kobylińska, Katarzyna, et al.
Veröffentlicht: (2023)
Make Interval Bound Propagation great again
von: Krukowski, Patryk, et al.
Veröffentlicht: (2024)
von: Krukowski, Patryk, et al.
Veröffentlicht: (2024)
Bounding Evidence and Estimating Log-Likelihood in VAE
von: Struski, Łukasz, et al.
Veröffentlicht: (2022)
von: Struski, Łukasz, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Self-Improving AI Agents through Self-Play
von: Chojecki, Przemyslaw
Veröffentlicht: (2025) -
The Geometry of Benchmarks: A New Path Toward AGI
von: Chojecki, Przemyslaw
Veröffentlicht: (2025) -
Psychometric Tests for AI Agents and Their Moduli Space
von: Chojecki, Przemyslaw
Veröffentlicht: (2025) -
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
von: Chojecki, Przemyslaw
Veröffentlicht: (2025) -
Model Science: getting serious about verification, explanation and control of AI systems
von: Biecek, Przemyslaw, et al.
Veröffentlicht: (2025)