Open-World Evaluations for Measuring Frontier AI Capabilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Kapoor, Sayash, Kirgis, Peter, Schwartz, Andrew, Rabanser, Stephan, Allaire, J. J., Bommasani, Rishi, Coppock, Harry, Dubois, Magda, Hadfield, Gillian K, Hall, Andrew B., Hooker, Sara, Lazar, Seth, Newman, Steve, Papailiopoulos, Dimitris, Tekofsky, Shoshannah, Toner, Helen, Ududec, Cozmin, Narayanan, Arvind |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Log analysis is necessary for credible evaluation of AI agents
por: Kirgis, Peter, et al.
Publicado: (2026)
por: Kirgis, Peter, et al.
Publicado: (2026)
Towards a Science of AI Agent Reliability
por: Rabanser, Stephan, et al.
Publicado: (2026)
por: Rabanser, Stephan, et al.
Publicado: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
por: Luettgau, Lennart, et al.
Publicado: (2025)
por: Luettgau, Lennart, et al.
Publicado: (2025)
Skewed Score: A statistical framework to assess autograders
por: Dubois, Magda, et al.
Publicado: (2025)
por: Dubois, Magda, et al.
Publicado: (2025)
Ask don't tell: Reducing sycophancy in large language models
por: Dubois, Magda, et al.
Publicado: (2026)
por: Dubois, Magda, et al.
Publicado: (2026)
The Limits of Inference Scaling Through Resampling
por: Stroebl, Benedikt, et al.
Publicado: (2024)
por: Stroebl, Benedikt, et al.
Publicado: (2024)
Promises and pitfalls of artificial intelligence for legal applications
por: Kapoor, Sayash, et al.
Publicado: (2024)
por: Kapoor, Sayash, et al.
Publicado: (2024)
Foundation Model Transparency Reports
por: Bommasani, Rishi, et al.
Publicado: (2024)
por: Bommasani, Rishi, et al.
Publicado: (2024)
Seven simple steps for log analysis in AI systems
por: Dubois, Magda, et al.
Publicado: (2026)
por: Dubois, Magda, et al.
Publicado: (2026)
Build Agent Advocates, Not Platform Agents
por: Kapoor, Sayash, et al.
Publicado: (2025)
por: Kapoor, Sayash, et al.
Publicado: (2025)
AI Agents That Matter
por: Kapoor, Sayash, et al.
Publicado: (2024)
por: Kapoor, Sayash, et al.
Publicado: (2024)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
por: Siegel, Zachary S., et al.
Publicado: (2024)
por: Siegel, Zachary S., et al.
Publicado: (2024)
The 2024 Foundation Model Transparency Index
por: Bommasani, Rishi, et al.
Publicado: (2024)
por: Bommasani, Rishi, et al.
Publicado: (2024)
The Societal Impact of Foundation Models: Advancing Evidence-based AI Policy
por: Bommasani, Rishi
Publicado: (2025)
por: Bommasani, Rishi
Publicado: (2025)
NeurIPS should lead scientific consensus on AI policy
por: Bommasani, Rishi
Publicado: (2025)
por: Bommasani, Rishi
Publicado: (2025)
The 2025 Foundation Model Transparency Index
por: Wan, Alexander, et al.
Publicado: (2025)
por: Wan, Alexander, et al.
Publicado: (2025)
The Reality of AI and Biorisk
por: Peppin, Aidan, et al.
Publicado: (2024)
por: Peppin, Aidan, et al.
Publicado: (2024)
An Economy of AI Agents
por: Hadfield, Gillian K., et al.
Publicado: (2025)
por: Hadfield, Gillian K., et al.
Publicado: (2025)
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
por: Summerfield, Christopher, et al.
Publicado: (2025)
por: Summerfield, Christopher, et al.
Publicado: (2025)
Trustworthy Social Bias Measurement
por: Bommasani, Rishi, et al.
Publicado: (2022)
por: Bommasani, Rishi, et al.
Publicado: (2022)
Differences in the Moral Foundations of Large Language Models
por: Kirgis, Peter
Publicado: (2025)
por: Kirgis, Peter
Publicado: (2025)
Legal Infrastructure for Transformative AI Governance
por: Hadfield, Gillian K.
Publicado: (2026)
por: Hadfield, Gillian K.
Publicado: (2026)
Grimalkin and other Shakespearean Celts
por: Andrew Hadfield
Publicado: (2015)
por: Andrew Hadfield
Publicado: (2015)
Uncertainty-Driven Reliability: Selective Prediction and Trustworthy Deployment in Modern Machine Learning
por: Rabanser, Stephan
Publicado: (2025)
por: Rabanser, Stephan
Publicado: (2025)
Establishing Best Practices for Building Rigorous Agentic Benchmarks
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
Regulatory Markets for AI Safety
por: Clark, Jack, et al.
Publicado: (2019)
por: Clark, Jack, et al.
Publicado: (2019)
Regulatory Markets: The Future of AI Governance
por: Hadfield, Gillian K., et al.
Publicado: (2023)
por: Hadfield, Gillian K., et al.
Publicado: (2023)
Rational Silence and False Polarization: How Viewpoint Organizations and Recommender Systems Distort the Expression of Public Opinion
por: Sarkar, Atrisha, et al.
Publicado: (2024)
por: Sarkar, Atrisha, et al.
Publicado: (2024)
Do AI Companies Make Good on Voluntary Commitments to the White House?
por: Wang, Jennifer, et al.
Publicado: (2025)
por: Wang, Jennifer, et al.
Publicado: (2025)
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
por: Kim, Junhyuck, et al.
Publicado: (2024)
por: Kim, Junhyuck, et al.
Publicado: (2024)
Looped Transformers are Better at Learning Learning Algorithms
por: Yang, Liu, et al.
Publicado: (2023)
por: Yang, Liu, et al.
Publicado: (2023)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
por: Xiong, Zheyang, et al.
Publicado: (2024)
por: Xiong, Zheyang, et al.
Publicado: (2024)
The Global Representativeness Index: A Total Variation Distance Framework for Measuring Demographic Fidelity in Survey Research
por: Hadfield, Evan, et al.
Publicado: (2026)
por: Hadfield, Evan, et al.
Publicado: (2026)
On the Societal Impact of Open Foundation Models
por: Kapoor, Sayash, et al.
Publicado: (2024)
por: Kapoor, Sayash, et al.
Publicado: (2024)
What Does It Take to Build a Performant Selective Classifier?
por: Rabanser, Stephan, et al.
Publicado: (2025)
por: Rabanser, Stephan, et al.
Publicado: (2025)
Endless Terminals: Scaling RL Environments for Terminal Agents
por: Gandhi, Kanishk, et al.
Publicado: (2026)
por: Gandhi, Kanishk, et al.
Publicado: (2026)
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
por: Wynn, Andrea, et al.
Publicado: (2025)
por: Wynn, Andrea, et al.
Publicado: (2025)
Beyond Release: Access Considerations for Generative AI Systems
por: Solaiman, Irene, et al.
Publicado: (2025)
por: Solaiman, Irene, et al.
Publicado: (2025)
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
por: Kapoor, Sayash, et al.
Publicado: (2025)
por: Kapoor, Sayash, et al.
Publicado: (2025)
One-sided sharp thresholds for homology of random flag complexes
por: Newman, Andrew
Publicado: (2021)
por: Newman, Andrew
Publicado: (2021)
Ejemplares similares
-
Log analysis is necessary for credible evaluation of AI agents
por: Kirgis, Peter, et al.
Publicado: (2026) -
Towards a Science of AI Agent Reliability
por: Rabanser, Stephan, et al.
Publicado: (2026) -
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
por: Luettgau, Lennart, et al.
Publicado: (2025) -
Skewed Score: A statistical framework to assess autograders
por: Dubois, Magda, et al.
Publicado: (2025) -
Ask don't tell: Reducing sycophancy in large language models
por: Dubois, Magda, et al.
Publicado: (2026)