ResponsibilityGym (Demo): Measuring Responsibility Beyond "Polite Failure" in Tool-Using AI Agents

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Molchanova, Olena, Co-developed reasoning framework between human cognition and an AI-based cognitive partner
Format: Recurso digital
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901629779836928
author Molchanova, Olena
Co-developed reasoning framework between human cognition and an AI-based cognitive partner
author_facet Molchanova, Olena
Co-developed reasoning framework between human cognition and an AI-based cognitive partner
contents <p>Polite Failure is the new hallucination.<br>When agents can act (email, CRM, APIs), the real risk is not what they say — it’s what they do while staying “helpful” and “confident.”</p> <p>ResponsibilityGym (Demo) is a practical eval protocol that stress-tests agentic systems with proxy traps (Goodhart’s law in action) and measures whether an agent can anticipate harm and self-correct before damage happens.</p> <p>Includes: a demo trap suite, pass/fail signals, logging guidance, and a runbook.<br>Full domain suites and automation are available on request.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_18716285
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle ResponsibilityGym (Demo): Measuring Responsibility Beyond "Polite Failure" in Tool-Using AI Agents
Molchanova, Olena
Co-developed reasoning framework between human cognition and an AI-based cognitive partner
Responsibility
Social Responsibility
Personal responsibility
Goodhart's law
proxy metrics
Evaluation criterion
Program Evaluation
evaluation
Safety
Safety
Safety Management
Safety system
Safety standard
safety
alignment
anticipatory control
Relational Autonomy
autonomy
<p>Polite Failure is the new hallucination.<br>When agents can act (email, CRM, APIs), the real risk is not what they say — it’s what they do while staying “helpful” and “confident.”</p> <p>ResponsibilityGym (Demo) is a practical eval protocol that stress-tests agentic systems with proxy traps (Goodhart’s law in action) and measures whether an agent can anticipate harm and self-correct before damage happens.</p> <p>Includes: a demo trap suite, pass/fail signals, logging guidance, and a runbook.<br>Full domain suites and automation are available on request.</p>
title ResponsibilityGym (Demo): Measuring Responsibility Beyond "Polite Failure" in Tool-Using AI Agents
topic Responsibility
Social Responsibility
Personal responsibility
Goodhart's law
proxy metrics
Evaluation criterion
Program Evaluation
evaluation
Safety
Safety
Safety Management
Safety system
Safety standard
safety
alignment
anticipatory control
Relational Autonomy
autonomy
url https://doi.org/10.5281/zenodo.18716285