Measuring What AI Systems Might Do: Towards A Measurement Science in AI
Fuente:
arXiv
Saved in:
| Main Authors: | Voudouris, Konstantinos, Thalmann, Mirko, Kipnis, Alex, Hernández-Orallo, José, Schulz, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025)
by: Rutar, Danaja, et al.
Published: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026)
by: Rabanser, Stephan, et al.
Published: (2026)
Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems
by: Gipiškis, Rokas, et al.
Published: (2024)
by: Gipiškis, Rokas, et al.
Published: (2024)
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
by: Kipnis, Alex, et al.
Published: (2024)
by: Kipnis, Alex, et al.
Published: (2024)
What should an AI assessor optimise for?
by: Romero-Alvarado, Daniel, et al.
Published: (2025)
by: Romero-Alvarado, Daniel, et al.
Published: (2025)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems
by: Buckley, Warren, et al.
Published: (2023)
by: Buckley, Warren, et al.
Published: (2023)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
Building, Reusing, and Generalizing Abstract Representations from Concrete Sequences
by: Wu, Shuchen, et al.
Published: (2024)
by: Wu, Shuchen, et al.
Published: (2024)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
PredictaBoard: Benchmarking LLM Score Predictability
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
by: Pacchiardi, Lorenzo, et al.
Published: (2025)
Towards Environmentally Equitable AI
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
by: Hajiesmaili, Mohammad, et al.
Published: (2024)
Regulating AI Adaptation: An Analysis of AI Medical Device Updates
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
Towards Socially and Environmentally Responsible AI
by: Li, Pengfei, et al.
Published: (2024)
by: Li, Pengfei, et al.
Published: (2024)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
Towards Effective Discrimination Testing for Generative AI
by: Zollo, Thomas P., et al.
Published: (2024)
by: Zollo, Thomas P., et al.
Published: (2024)
Defining AI Models and AI Systems: A Framework to Resolve the Boundary Problem
by: Sun, Yuanyuan, et al.
Published: (2026)
by: Sun, Yuanyuan, et al.
Published: (2026)
Why am I Still Seeing This: Measuring the Effectiveness Of Ad Controls and Explanations in AI-Mediated Ad Targeting Systems
by: Castleman, Jane, et al.
Published: (2024)
by: Castleman, Jane, et al.
Published: (2024)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
by: Testini, Irene, et al.
Published: (2025)
by: Testini, Irene, et al.
Published: (2025)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?
by: Dreksler, Noemi, et al.
Published: (2025)
by: Dreksler, Noemi, et al.
Published: (2025)
FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks
by: Wang, Miles, et al.
Published: (2026)
by: Wang, Miles, et al.
Published: (2026)
What happens when generative AI models train recursively on each others' outputs?
by: Vu, Hung Anh, et al.
Published: (2025)
by: Vu, Hung Anh, et al.
Published: (2025)
The AI Companion in Education: Analyzing the Pedagogical Potential of ChatGPT in Computer Science and Engineering
by: He, Zhangying, et al.
Published: (2024)
by: He, Zhangying, et al.
Published: (2024)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
Being Considerate as a Pathway Towards Pluralistic Alignment for Agentic AI
by: Alamdari, Parand A., et al.
Published: (2024)
by: Alamdari, Parand A., et al.
Published: (2024)
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
by: Jurenka, Irina, et al.
Published: (2024)
by: Jurenka, Irina, et al.
Published: (2024)
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
by: Qureshi, Taro, et al.
Published: (2026)
by: Qureshi, Taro, et al.
Published: (2026)
Thousands of AI Authors on the Future of AI
by: Grace, Katja, et al.
Published: (2024)
by: Grace, Katja, et al.
Published: (2024)
Generative AI Meets Future Cities: Towards an Era of Autonomous Urban Intelligence
by: Wang, Dongjie, et al.
Published: (2023)
by: Wang, Dongjie, et al.
Published: (2023)
AI Data Development: A Scorecard for the System Card Framework
by: Bahiru, Tadesse K., et al.
Published: (2025)
by: Bahiru, Tadesse K., et al.
Published: (2025)
Safe and Certifiable AI Systems: Concepts, Challenges, and Lessons Learned
by: Schweighofer, Kajetan, et al.
Published: (2025)
by: Schweighofer, Kajetan, et al.
Published: (2025)
Explainable AI Systems Must Be Contestable: Here's How to Make It Happen
by: Moreira, Catarina, et al.
Published: (2025)
by: Moreira, Catarina, et al.
Published: (2025)
Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society
by: Hartmann, David, et al.
Published: (2024)
by: Hartmann, David, et al.
Published: (2024)
What constitutes a Deep Fake? The blurry line between legitimate processing and manipulation under the EU AI Act
by: Meding, Kristof, et al.
Published: (2024)
by: Meding, Kristof, et al.
Published: (2024)
AI Toolkit: Libraries and Essays for Exploring the Technology and Ethics of AI
by: Ho, Levin, et al.
Published: (2025)
by: Ho, Levin, et al.
Published: (2025)
Lessons for Editors of AI Incidents from the AI Incident Database
by: Paeth, Kevin, et al.
Published: (2024)
by: Paeth, Kevin, et al.
Published: (2024)
Mapping the Potential of Explainable AI for Fairness Along the AI Lifecycle
by: Deck, Luca, et al.
Published: (2024)
by: Deck, Luca, et al.
Published: (2024)
Similar Items
-
Cognitive Science-Inspired Evaluation of Core Capabilities for Object Understanding in AI
by: Rutar, Danaja, et al.
Published: (2025) -
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024) -
Towards a Science of AI Agent Reliability
by: Rabanser, Stephan, et al.
Published: (2026) -
Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems
by: Gipiškis, Rokas, et al.
Published: (2024) -
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
by: Kipnis, Alex, et al.
Published: (2024)