Agent Benchmarks Fail Public Sector Requirements
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rystrøm, Jonathan, Schmitz, Chris, Korgul, Karolina, Batzner, Jan, Russell, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Oversight Structures for Agentic AI in Public-Sector Organizations
von: Schmitz, Chris, et al.
Veröffentlicht: (2025)
von: Schmitz, Chris, et al.
Veröffentlicht: (2025)
Grounding Text Embeddings in Stakeholder Associations
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2026)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2026)
OxEnsemble: Fair Ensembles for Low-Data Classification
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
von: Korgul, Karolina, et al.
Veröffentlicht: (2025)
von: Korgul, Karolina, et al.
Veröffentlicht: (2025)
Exposing Assumptions in AI Benchmarks through Cognitive Modelling
von: Rystrøm, Jonathan H., et al.
Veröffentlicht: (2024)
von: Rystrøm, Jonathan H., et al.
Veröffentlicht: (2024)
A Systems Thinking Approach to Algorithmic Fairness
von: Lam, Chris
Veröffentlicht: (2024)
von: Lam, Chris
Veröffentlicht: (2024)
Deepfakes on Demand: the rise of accessible non-consensual deepfake image generators
von: Hawkins, Will, et al.
Veröffentlicht: (2025)
von: Hawkins, Will, et al.
Veröffentlicht: (2025)
Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2025)
von: Shen, Judy Hanwen, et al.
Veröffentlicht: (2025)
Perceptions of AI Across Sectors: A Comparative Review of Public Attitudes
von: Bialy, Filip, et al.
Veröffentlicht: (2025)
von: Bialy, Filip, et al.
Veröffentlicht: (2025)
The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
von: Bent, Brinnae
Veröffentlicht: (2025)
von: Bent, Brinnae
Veröffentlicht: (2025)
OxonFair: A Flexible Toolkit for Algorithmic Fairness
von: Delaney, Eoin, et al.
Veröffentlicht: (2024)
von: Delaney, Eoin, et al.
Veröffentlicht: (2024)
Algorithmic Administration and the EU AI Act: Legal Principles for Public Sector Use of AI
von: Pavlidis, Georgios, et al.
Veröffentlicht: (2026)
von: Pavlidis, Georgios, et al.
Veröffentlicht: (2026)
Responsible AI Adoption in the Public Sector: A Data-Centric Taxonomy of AI Adoption Challenges
von: Nikiforova, Anastasija, et al.
Veröffentlicht: (2025)
von: Nikiforova, Anastasija, et al.
Veröffentlicht: (2025)
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
von: Lee, Kyeongryul, et al.
Veröffentlicht: (2025)
von: Lee, Kyeongryul, et al.
Veröffentlicht: (2025)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
von: Li, Miles Q., et al.
Veröffentlicht: (2026)
Failing on Bias Mitigation: A Case Study on the Challenges of Fairness in Government Data
von: Bo, Hongbo, et al.
Veröffentlicht: (2026)
von: Bo, Hongbo, et al.
Veröffentlicht: (2026)
No Transfers Required: Integrating Last Mile with Public Transit Using Opti-Mile
von: Altaf, Raashid, et al.
Veröffentlicht: (2023)
von: Altaf, Raashid, et al.
Veröffentlicht: (2023)
Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight
von: Chappidi, Shreya, et al.
Veröffentlicht: (2025)
von: Chappidi, Shreya, et al.
Veröffentlicht: (2025)
Technical Requirements for Halting Dangerous AI Activities
von: Barnett, Peter, et al.
Veröffentlicht: (2025)
von: Barnett, Peter, et al.
Veröffentlicht: (2025)
E-LENS: User Requirements-Oriented AI Ethics Assurance
von: Zhou, Jianlong, et al.
Veröffentlicht: (2025)
von: Zhou, Jianlong, et al.
Veröffentlicht: (2025)
When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies
von: Popchanovska, Evgenija, et al.
Veröffentlicht: (2026)
von: Popchanovska, Evgenija, et al.
Veröffentlicht: (2026)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
Trademark Search, Artificial Intelligence and the Role of the Private Sector
von: Katyal, Sonia, et al.
Veröffentlicht: (2026)
von: Katyal, Sonia, et al.
Veröffentlicht: (2026)
AI-Mediated Communication Can Steer Collective Opinion
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2026)
von: Tsirtsis, Stratis, et al.
Veröffentlicht: (2026)
Towards an AI Observatory for the Nuclear Sector: A tool for anticipatory governance
von: Verma, Aditi, et al.
Veröffentlicht: (2025)
von: Verma, Aditi, et al.
Veröffentlicht: (2025)
Initial results of the Digital Consciousness Model
von: Shiller, Derek, et al.
Veröffentlicht: (2026)
von: Shiller, Derek, et al.
Veröffentlicht: (2026)
Agents of Chaos
von: Shapira, Natalie, et al.
Veröffentlicht: (2026)
von: Shapira, Natalie, et al.
Veröffentlicht: (2026)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
Assessing Model-Agnostic XAI Methods against EU AI Act Explainability Requirements
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
How should AI decisions be explained? Requirements for Explanations from the Perspective of European Law
von: Fresz, Benjamin, et al.
Veröffentlicht: (2024)
von: Fresz, Benjamin, et al.
Veröffentlicht: (2024)
The Human Condition as Reflected in Contemporary Large Language Models
von: Neuman, W. Russell
Veröffentlicht: (2026)
von: Neuman, W. Russell
Veröffentlicht: (2026)
An Ontology for Representing Curriculum and Learning Material
von: Christou, Antrea, et al.
Veröffentlicht: (2025)
von: Christou, Antrea, et al.
Veröffentlicht: (2025)
AI and the Future of Digital Public Squares
von: Goldberg, Beth, et al.
Veröffentlicht: (2024)
von: Goldberg, Beth, et al.
Veröffentlicht: (2024)
Education in the Era of Neurosymbolic AI
von: Jaldi, Chris Davis, et al.
Veröffentlicht: (2024)
von: Jaldi, Chris Davis, et al.
Veröffentlicht: (2024)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
von: Zou, Andy, et al.
Veröffentlicht: (2025)
von: Zou, Andy, et al.
Veröffentlicht: (2025)
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
von: McCaslin, Tegan, et al.
Veröffentlicht: (2025)
von: McCaslin, Tegan, et al.
Veröffentlicht: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
von: Bowen, Dillon, et al.
Veröffentlicht: (2025)
Simulating Society Requires Simulating Thought
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Chance Jiajie, et al.
Veröffentlicht: (2025)
Analyzing the Ethical Logic of Six Large Language Models
von: Neuman, W. Russell, et al.
Veröffentlicht: (2025)
von: Neuman, W. Russell, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Oversight Structures for Agentic AI in Public-Sector Organizations
von: Schmitz, Chris, et al.
Veröffentlicht: (2025) -
Grounding Text Embeddings in Stakeholder Associations
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2026) -
OxEnsemble: Fair Ensembles for Low-Data Classification
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025) -
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025) -
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
von: Korgul, Karolina, et al.
Veröffentlicht: (2025)