The Geometry of Benchmarks: A New Path Toward AGI
Fuente:
arXiv
Saved in:
| Main Author: | Chojecki, Przemyslaw |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Psychometric Tests for AI Agents and Their Moduli Space
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Mathematics and Coding are Universal AI Benchmarks
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
On the Geometry of Receiver Operating Characteristic and Precision-Recall Curves
by: Sameni, Reza
Published: (2025)
by: Sameni, Reza
Published: (2025)
Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry Adaptive
by: Farghly, Tyler, et al.
Published: (2025)
by: Farghly, Tyler, et al.
Published: (2025)
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
by: Yu, Hao
Published: (2025)
by: Yu, Hao
Published: (2025)
Towards Bayesian Data Selection
by: Rodemann, Julian
Published: (2024)
by: Rodemann, Julian
Published: (2024)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Statistical inference with belief functions: A survey
by: Cuzzolin, Fabio
Published: (2026)
by: Cuzzolin, Fabio
Published: (2026)
A Quantitative Characterization of Forgetting in Post-Training
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
Deep Ensembles for Epistemic Uncertainty: A Frequentist Perspective
by: Jain, Anchit, et al.
Published: (2025)
by: Jain, Anchit, et al.
Published: (2025)
A Fine-Grained Understanding of Uniform Convergence for Halfspaces
by: Kontorovich, Aryeh, et al.
Published: (2026)
by: Kontorovich, Aryeh, et al.
Published: (2026)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees
by: Yan, Hedong
Published: (2025)
by: Yan, Hedong
Published: (2025)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
by: Hazard, Christopher J., et al.
Published: (2025)
by: Hazard, Christopher J., et al.
Published: (2025)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
A note on the impossibility of conditional PAC-efficient reasoning in large language models
by: Zeng, Hao
Published: (2025)
by: Zeng, Hao
Published: (2025)
A Statistical Analysis of Deep Federated Learning for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2024)
by: Chakraborty, Saptarshi, et al.
Published: (2024)
Low-Dimensional Adaptation of Rectified Flow: A Diffusion and Stochastic Localization Perspective
by: Roy, Saptarshi, et al.
Published: (2026)
by: Roy, Saptarshi, et al.
Published: (2026)
A comparative study of conformal prediction methods for valid uncertainty quantification in machine learning
by: Dewolf, Nicolas
Published: (2024)
by: Dewolf, Nicolas
Published: (2024)
A Unified Pair-GRPO Family: From Implicit to Explicit Preference Constraints for Stable and General RL Alignment
by: Yu, Hao
Published: (2026)
by: Yu, Hao
Published: (2026)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Efficient Knowledge Distillation via Curriculum Extraction
by: Gupta, Shivam, et al.
Published: (2025)
by: Gupta, Shivam, et al.
Published: (2025)
Risk Analysis and Design Against Adversarial Actions
by: Campi, Marco C., et al.
Published: (2025)
by: Campi, Marco C., et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Cross-regularization: Adaptive Model Complexity through Validation Gradients
by: Brito, Carlos Stein
Published: (2025)
by: Brito, Carlos Stein
Published: (2025)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
by: Xu, Yichen, et al.
Published: (2025)
by: Xu, Yichen, et al.
Published: (2025)
Dense associative memory for Gaussian distributions
by: Tankala, Chandan, et al.
Published: (2025)
by: Tankala, Chandan, et al.
Published: (2025)
On the Statistical Capacity of Deep Generative Models
by: Tam, Edric, et al.
Published: (2025)
by: Tam, Edric, et al.
Published: (2025)
Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust Optimization
by: Liu, Shuang, et al.
Published: (2025)
by: Liu, Shuang, et al.
Published: (2025)
When Can We Reuse a Calibration Set for Multiple Conformal Predictions?
by: Balinsky, A. A., et al.
Published: (2025)
by: Balinsky, A. A., et al.
Published: (2025)
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
by: Liu, Hanwen, et al.
Published: (2025)
by: Liu, Hanwen, et al.
Published: (2025)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
What is causal about causal models and representations?
by: Jørgensen, Frederik Hytting, et al.
Published: (2025)
by: Jørgensen, Frederik Hytting, et al.
Published: (2025)
Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability
by: Yu, Lijia, et al.
Published: (2025)
by: Yu, Lijia, et al.
Published: (2025)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
by: Baharav, Tavor Z., et al.
Published: (2025)
by: Baharav, Tavor Z., et al.
Published: (2025)
Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
by: Shen, Zhaiming, et al.
Published: (2025)
by: Shen, Zhaiming, et al.
Published: (2025)
Conformal Prediction for Privacy-Preserving Machine Learning
by: Balinsky, Alexander David, et al.
Published: (2025)
by: Balinsky, Alexander David, et al.
Published: (2025)
Similar Items
-
Psychometric Tests for AI Agents and Their Moduli Space
by: Chojecki, Przemyslaw
Published: (2025) -
Mathematics and Coding are Universal AI Benchmarks
by: Chojecki, Przemyslaw
Published: (2025) -
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
by: Chojecki, Przemyslaw
Published: (2025) -
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025) -
On the Geometry of Receiver Operating Characteristic and Precision-Recall Curves
by: Sameni, Reza
Published: (2025)