Position: AI Evaluations Should be Grounded on a Theory of Capability
Fuente:
arXiv
Saved in:
| Main Authors: | Jo, Nathanael, Wilson, Ashia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025)
by: Bowen, Dillon, et al.
Published: (2025)
DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
by: De Simone, Zoe, et al.
Published: (2023)
by: De Simone, Zoe, et al.
Published: (2023)
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
by: Tang, Zeyu, et al.
Published: (2025)
by: Tang, Zeyu, et al.
Published: (2025)
Fairness Evaluation for Uplift Modeling in the Absence of Ground Truth
by: Kadioglu, Serdar, et al.
Published: (2024)
by: Kadioglu, Serdar, et al.
Published: (2024)
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025)
by: Kim, HyunJin, et al.
Published: (2025)
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
by: Vishwarupe, Varad, et al.
Published: (2026)
by: Vishwarupe, Varad, et al.
Published: (2026)
Incentives shape how humans co-create with generative AI
by: Jo, Nathanael, et al.
Published: (2026)
by: Jo, Nathanael, et al.
Published: (2026)
Evaluating AI Group Fairness: a Fuzzy Logic Perspective
by: Krasanakis, Emmanouil, et al.
Published: (2024)
by: Krasanakis, Emmanouil, et al.
Published: (2024)
Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data
by: Nolte, Henrik, et al.
Published: (2025)
by: Nolte, Henrik, et al.
Published: (2025)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
Watermarking Should Be Treated as a Monitoring Primitive
by: Aremu, Toluwani, et al.
Published: (2026)
by: Aremu, Toluwani, et al.
Published: (2026)
Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
by: Reuel, Anka, et al.
Published: (2025)
by: Reuel, Anka, et al.
Published: (2025)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Position Paper: If Innovation in AI Systematically Violates Fundamental Rights, Is It Innovation at All?
by: Castañeira, Josu Eguiluz, et al.
Published: (2025)
by: Castañeira, Josu Eguiluz, et al.
Published: (2025)
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
by: Jurenka, Irina, et al.
Published: (2024)
by: Jurenka, Irina, et al.
Published: (2024)
Synthetic Data and the Shifting Ground of Truth
by: Offenhuber, Dietmar
Published: (2025)
by: Offenhuber, Dietmar
Published: (2025)
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
by: Patwardhan, Tejal, et al.
Published: (2025)
by: Patwardhan, Tejal, et al.
Published: (2025)
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
by: Wang, Miles, et al.
Published: (2026)
by: Wang, Miles, et al.
Published: (2026)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models I: The Task-Query Architecture
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Alignment has a Fantasia Problem
by: Jo, Nathanael, et al.
Published: (2026)
by: Jo, Nathanael, et al.
Published: (2026)
The Approximate Fisher Influence Function: Faster Estimation of Data Influence in Statistical Models
by: Lev, Omri, et al.
Published: (2024)
by: Lev, Omri, et al.
Published: (2024)
AssurAI: Experience with Constructing Korean Socio-cultural Datasets to Discover Potential Risks of Generative AI
by: Lim, Chae-Gyun, et al.
Published: (2025)
by: Lim, Chae-Gyun, et al.
Published: (2025)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Position: Cracking the Code of Cascading Disparity Towards Marginalized Communities
by: Farnadi, Golnoosh, et al.
Published: (2024)
by: Farnadi, Golnoosh, et al.
Published: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
by: Zhang, Andy K., et al.
Published: (2024)
by: Zhang, Andy K., et al.
Published: (2024)
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
by: Tariq, Amara, et al.
Published: (2025)
by: Tariq, Amara, et al.
Published: (2025)
Thousands of AI Authors on the Future of AI
by: Grace, Katja, et al.
Published: (2024)
by: Grace, Katja, et al.
Published: (2024)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
Position: Capability Control Should be a Separate Goal From Alignment
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
by: Siddiqui, Shoaib Ahmed, et al.
Published: (2026)
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026)
by: Rabanser, Stephan, et al.
Published: (2026)
Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants
by: Ding, Shi, et al.
Published: (2025)
by: Ding, Shi, et al.
Published: (2025)
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
by: Qureshi, Taro, et al.
Published: (2026)
by: Qureshi, Taro, et al.
Published: (2026)
AI Toolkit: Libraries and Essays for Exploring the Technology and Ethics of AI
by: Ho, Levin, et al.
Published: (2025)
by: Ho, Levin, et al.
Published: (2025)
Lessons for Editors of AI Incidents from the AI Incident Database
by: Paeth, Kevin, et al.
Published: (2024)
by: Paeth, Kevin, et al.
Published: (2024)
Regulating AI Adaptation: An Analysis of AI Medical Device Updates
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
Mapping the Potential of Explainable AI for Fairness Along the AI Lifecycle
by: Deck, Luca, et al.
Published: (2024)
by: Deck, Luca, et al.
Published: (2024)
A Consequentialist Critique of Binary Classification Evaluation: Theory, Practice, and Tools
by: Flores, Gerardo, et al.
Published: (2025)
by: Flores, Gerardo, et al.
Published: (2025)
Biothreat Benchmark Generation Framework for Evaluating Frontier AI Models III: Implementing the Bacterial Biothreat Benchmark (B3) Dataset
by: Ackerman, Gary, et al.
Published: (2025)
by: Ackerman, Gary, et al.
Published: (2025)
Similar Items
-
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
by: Bowen, Dillon, et al.
Published: (2025) -
DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
by: De Simone, Zoe, et al.
Published: (2023) -
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
by: Schaeffer, Rylan, et al.
Published: (2025) -
Position: Beyond Sensitive Attributes, ML Fairness Should Quantify Structural Injustice via Social Determinants
by: Tang, Zeyu, et al.
Published: (2025) -
Fairness Evaluation for Uplift Modeling in the Absence of Ground Truth
by: Kadioglu, Serdar, et al.
Published: (2024)