Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Eriksson, Maria, Purificato, Erasmo, Noroozian, Arman, Vinagre, Joao, Chaslot, Guillaume, Gomez, Emilia, Fernandez-Llorca, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI
di: de Cerqueira, José Antonio Siqueira, et al.
Pubblicazione: (2024)
di: de Cerqueira, José Antonio Siqueira, et al.
Pubblicazione: (2024)
Generative AI and the Future of the Digital Commons: Five Open Questions and Knowledge Gaps
di: Noroozian, Arman, et al.
Pubblicazione: (2025)
di: Noroozian, Arman, et al.
Pubblicazione: (2025)
Trusting AI in High-stake Decision Making
di: Saffarini, Ali
Pubblicazione: (2023)
di: Saffarini, Ali
Pubblicazione: (2023)
Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
di: Sahai, Sattvik, et al.
Pubblicazione: (2025)
di: Sahai, Sattvik, et al.
Pubblicazione: (2025)
Seeing The Words: Evaluating AI-generated Biblical Art
di: Makimei, Hidde, et al.
Pubblicazione: (2025)
di: Makimei, Hidde, et al.
Pubblicazione: (2025)
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
di: Pant, Aakash, et al.
Pubblicazione: (2026)
di: Pant, Aakash, et al.
Pubblicazione: (2026)
Towards Friendly AI: A Comprehensive Review and New Perspectives on Human-AI Alignment
di: Sun, Qiyang, et al.
Pubblicazione: (2024)
di: Sun, Qiyang, et al.
Pubblicazione: (2024)
The Station: An Open-World Environment for AI-Driven Discovery
di: Chung, Stephen, et al.
Pubblicazione: (2025)
di: Chung, Stephen, et al.
Pubblicazione: (2025)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
di: Altamirano-Gomez, Gerardo, et al.
Pubblicazione: (2023)
di: Altamirano-Gomez, Gerardo, et al.
Pubblicazione: (2023)
Threats and Opportunities in AI-generated Images for Armed Forces
di: Meier, Raphael
Pubblicazione: (2025)
di: Meier, Raphael
Pubblicazione: (2025)
Making High-Level AI Design Decisions Explicit Using a Binary Stream System-Designation Approach
di: Mossbridge, Julia
Pubblicazione: (2024)
di: Mossbridge, Julia
Pubblicazione: (2024)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
Diverse AI Personas Can Mitigate the Homogenization Effect in Human-AI Collaborative Ideation
di: Wan, Yun, et al.
Pubblicazione: (2025)
di: Wan, Yun, et al.
Pubblicazione: (2025)
FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare
di: Lekadir, Karim, et al.
Pubblicazione: (2023)
di: Lekadir, Karim, et al.
Pubblicazione: (2023)
The Landscape of Generative AI in Information Systems: A Synthesis of Secondary Reviews and Research Agendas
di: Jarzębowicz, Aleksander, et al.
Pubblicazione: (2026)
di: Jarzębowicz, Aleksander, et al.
Pubblicazione: (2026)
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
di: Mathys, Melanie, et al.
Pubblicazione: (2024)
di: Mathys, Melanie, et al.
Pubblicazione: (2024)
Antagonistic AI
di: Cai, Alice, et al.
Pubblicazione: (2024)
di: Cai, Alice, et al.
Pubblicazione: (2024)
A Framework for Responsible AI Systems: Building Societal Trust through Domain Definition, Trustworthy AI Design, Auditability, Accountability, and Governance
di: Herrera-Poyatos, Andrés, et al.
Pubblicazione: (2025)
di: Herrera-Poyatos, Andrés, et al.
Pubblicazione: (2025)
Evaluating Large Language Models for Causal Modeling
di: Razouk, Houssam, et al.
Pubblicazione: (2024)
di: Razouk, Houssam, et al.
Pubblicazione: (2024)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
di: Schaeffer, Joachim, et al.
Pubblicazione: (2026)
di: Schaeffer, Joachim, et al.
Pubblicazione: (2026)
We Need a New Ethics for a World of AI Agents
di: Gabriel, Iason, et al.
Pubblicazione: (2025)
di: Gabriel, Iason, et al.
Pubblicazione: (2025)
Moral Responsibility or Obedience: What Do We Want from AI?
di: Boland, Joseph
Pubblicazione: (2025)
di: Boland, Joseph
Pubblicazione: (2025)
Data Feminism for AI
di: Klein, Lauren, et al.
Pubblicazione: (2024)
di: Klein, Lauren, et al.
Pubblicazione: (2024)
Developing trustworthy AI applications with foundation models
di: Mock, Michael, et al.
Pubblicazione: (2024)
di: Mock, Michael, et al.
Pubblicazione: (2024)
ETOM: A Five-Level Benchmark for Evaluating Tool Orchestration within the MCP Ecosystem
di: Dong, Jia-Kai, et al.
Pubblicazione: (2025)
di: Dong, Jia-Kai, et al.
Pubblicazione: (2025)
Right-to-Act: A Pre-Execution Non-Compensatory Decision Protocol for AI Systems
di: Lavi, Gadi
Pubblicazione: (2026)
di: Lavi, Gadi
Pubblicazione: (2026)
AI for All: Identifying AI incidents Related to Diversity and Inclusion
di: Shams, Rifat Ara, et al.
Pubblicazione: (2024)
di: Shams, Rifat Ara, et al.
Pubblicazione: (2024)
Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education
di: Hooshyar, Danial, et al.
Pubblicazione: (2023)
di: Hooshyar, Danial, et al.
Pubblicazione: (2023)
AI Literacy in K-12 and Higher Education in the Wake of Generative AI: An Integrative Review
di: Gu, Xingjian, et al.
Pubblicazione: (2025)
di: Gu, Xingjian, et al.
Pubblicazione: (2025)
Event-based Solutions for Human-centered Applications: A Comprehensive Review
di: Adra, Mira, et al.
Pubblicazione: (2025)
di: Adra, Mira, et al.
Pubblicazione: (2025)
What Does 'Human-Centred AI' Mean?
di: Guest, Olivia
Pubblicazione: (2025)
di: Guest, Olivia
Pubblicazione: (2025)
Provocations from the Humanities for Generative AI Research
di: Klein, Lauren, et al.
Pubblicazione: (2025)
di: Klein, Lauren, et al.
Pubblicazione: (2025)
SIDEs: Separating Idealization from Deceptive Explanations in xAI
di: Sullivan, Emily
Pubblicazione: (2024)
di: Sullivan, Emily
Pubblicazione: (2024)
CURATe: Benchmarking Personalised Alignment of Conversational AI Assistants
di: Alberts, Lize, et al.
Pubblicazione: (2024)
di: Alberts, Lize, et al.
Pubblicazione: (2024)
Fanar: An Arabic-Centric Multimodal Generative AI Platform
di: Fanar Team, et al.
Pubblicazione: (2025)
di: Fanar Team, et al.
Pubblicazione: (2025)
Reward is not enough: can we liberate AI from the reinforcement learning paradigm?
di: Glukhov, Vacslav
Pubblicazione: (2022)
di: Glukhov, Vacslav
Pubblicazione: (2022)
A Review of Pseudo-Labeling for Computer Vision
di: Kage, Patrick, et al.
Pubblicazione: (2024)
di: Kage, Patrick, et al.
Pubblicazione: (2024)
Rip Current Segmentation: A Novel Benchmark and YOLOv8 Baseline Results
di: Dumitriu, Andrei, et al.
Pubblicazione: (2025)
di: Dumitriu, Andrei, et al.
Pubblicazione: (2025)
Evaluating AI-Enabled deception vulnerability amongst Sub-Saharan-Africa migrants
di: Oluwasanya, Deborah
Pubblicazione: (2026)
di: Oluwasanya, Deborah
Pubblicazione: (2026)
AI and the Problem of Knowledge Collapse
di: Peterson, Andrew J.
Pubblicazione: (2024)
di: Peterson, Andrew J.
Pubblicazione: (2024)
Documenti analoghi
-
Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI
di: de Cerqueira, José Antonio Siqueira, et al.
Pubblicazione: (2024) -
Generative AI and the Future of the Digital Commons: Five Open Questions and Knowledge Gaps
di: Noroozian, Arman, et al.
Pubblicazione: (2025) -
Trusting AI in High-stake Decision Making
di: Saffarini, Ali
Pubblicazione: (2023) -
Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
di: Sahai, Sattvik, et al.
Pubblicazione: (2025) -
Seeing The Words: Evaluating AI-generated Biblical Art
di: Makimei, Hidde, et al.
Pubblicazione: (2025)