Psychometric Tests for AI Agents and Their Moduli Space
Fuente:
arXiv
Saved in:
| Main Author: | Chojecki, Przemyslaw |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Mathematics and Coding are Universal AI Benchmarks
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Enhancing Conformal Prediction Using E-Test Statistics
by: Balinsky, A. A., et al.
Published: (2024)
by: Balinsky, A. A., et al.
Published: (2024)
Cost-optimal Sequential Testing via Doubly Robust Q-learning
by: Zhou, Doudou, et al.
Published: (2026)
by: Zhou, Doudou, et al.
Published: (2026)
A Computational Theory for Efficient Mini Agent Evaluation with Causal Guarantees
by: Yan, Hedong
Published: (2025)
by: Yan, Hedong
Published: (2025)
Adaptive auditing of AI systems with anytime-valid guarantees
by: Zhou, Siyu, et al.
Published: (2026)
by: Zhou, Siyu, et al.
Published: (2026)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Interaction Testing in Variation Analysis
by: Plecko, Drago
Published: (2024)
by: Plecko, Drago
Published: (2024)
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
by: Dong, Zihan, et al.
Published: (2026)
by: Dong, Zihan, et al.
Published: (2026)
An Operational Kardashev-Style Scale for Autonomous AI - Towards AGI and Superintelligence
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Efficient Knowledge Distillation via Curriculum Extraction
by: Gupta, Shivam, et al.
Published: (2025)
by: Gupta, Shivam, et al.
Published: (2025)
Risk Analysis and Design Against Adversarial Actions
by: Campi, Marco C., et al.
Published: (2025)
by: Campi, Marco C., et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
by: Hazard, Christopher J., et al.
Published: (2025)
by: Hazard, Christopher J., et al.
Published: (2025)
Cross-regularization: Adaptive Model Complexity through Validation Gradients
by: Brito, Carlos Stein
Published: (2025)
by: Brito, Carlos Stein
Published: (2025)
Residual Feature Integration is Sufficient to Prevent Negative Transfer
by: Xu, Yichen, et al.
Published: (2025)
by: Xu, Yichen, et al.
Published: (2025)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Dense associative memory for Gaussian distributions
by: Tankala, Chandan, et al.
Published: (2025)
by: Tankala, Chandan, et al.
Published: (2025)
On the Statistical Capacity of Deep Generative Models
by: Tam, Edric, et al.
Published: (2025)
by: Tam, Edric, et al.
Published: (2025)
Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust Optimization
by: Liu, Shuang, et al.
Published: (2025)
by: Liu, Shuang, et al.
Published: (2025)
When Can We Reuse a Calibration Set for Multiple Conformal Predictions?
by: Balinsky, A. A., et al.
Published: (2025)
by: Balinsky, A. A., et al.
Published: (2025)
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
by: Liu, Hanwen, et al.
Published: (2025)
by: Liu, Hanwen, et al.
Published: (2025)
On the Provable Performance Guarantee of Efficient Reasoning Models
by: Zeng, Hao, et al.
Published: (2025)
by: Zeng, Hao, et al.
Published: (2025)
What is causal about causal models and representations?
by: Jørgensen, Frederik Hytting, et al.
Published: (2025)
by: Jørgensen, Frederik Hytting, et al.
Published: (2025)
Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability
by: Yu, Lijia, et al.
Published: (2025)
by: Yu, Lijia, et al.
Published: (2025)
The Good, the Bad, and the Sampled: a No-Regret Approach to Safe Online Classification
by: Baharav, Tavor Z., et al.
Published: (2025)
by: Baharav, Tavor Z., et al.
Published: (2025)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
On the Geometry of Receiver Operating Characteristic and Precision-Recall Curves
by: Sameni, Reza
Published: (2025)
by: Sameni, Reza
Published: (2025)
Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
by: Shen, Zhaiming, et al.
Published: (2025)
by: Shen, Zhaiming, et al.
Published: (2025)
Conformal Prediction for Privacy-Preserving Machine Learning
by: Balinsky, Alexander David, et al.
Published: (2025)
by: Balinsky, Alexander David, et al.
Published: (2025)
Sample Complexity of Bias Detection with Subsampled Point-to-Subspace Distances
by: Matilla, German Martinez, et al.
Published: (2025)
by: Matilla, German Martinez, et al.
Published: (2025)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Convergence of Shallow ReLU Networks on Weakly Interacting Data
by: Dana, Léo, et al.
Published: (2025)
by: Dana, Léo, et al.
Published: (2025)
Foundations of Top-$k$ Decoding For Language Models
by: Noarov, Georgy, et al.
Published: (2025)
by: Noarov, Georgy, et al.
Published: (2025)
Cyclic Counterfactuals under Shift-Scale Interventions
by: Saha, Saptarshi, et al.
Published: (2025)
by: Saha, Saptarshi, et al.
Published: (2025)
Learning with Differentially Private (Sliced) Wasserstein Gradients
by: Rodríguez-Vítores, David, et al.
Published: (2025)
by: Rodríguez-Vítores, David, et al.
Published: (2025)
A note on the impossibility of conditional PAC-efficient reasoning in large language models
by: Zeng, Hao
Published: (2025)
by: Zeng, Hao
Published: (2025)
Deep Ensembles for Epistemic Uncertainty: A Frequentist Perspective
by: Jain, Anchit, et al.
Published: (2025)
by: Jain, Anchit, et al.
Published: (2025)
Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry Adaptive
by: Farghly, Tyler, et al.
Published: (2025)
by: Farghly, Tyler, et al.
Published: (2025)
Similar Items
-
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025) -
Self-Improving AI Agents through Self-Play
by: Chojecki, Przemyslaw
Published: (2025) -
Mathematics and Coding are Universal AI Benchmarks
by: Chojecki, Przemyslaw
Published: (2025) -
Enhancing Conformal Prediction Using E-Test Statistics
by: Balinsky, A. A., et al.
Published: (2024) -
Cost-optimal Sequential Testing via Doubly Robust Q-learning
by: Zhou, Doudou, et al.
Published: (2026)