ResponsibilityGym (Demo): Measuring Responsibility Beyond "Polite Failure" in Tool-Using AI Agents
Fuente:
Zenodo
Gespeichert in:
| Hauptverfasser: | , |
|---|---|
| Format: | Recurso digital |
| Veröffentlicht: |
Zenodo
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866901629779836928 |
|---|---|
| author | Molchanova, Olena Co-developed reasoning framework between human cognition and an AI-based cognitive partner |
| author_facet | Molchanova, Olena Co-developed reasoning framework between human cognition and an AI-based cognitive partner |
| contents | <p>Polite Failure is the new hallucination.<br>When agents can act (email, CRM, APIs), the real risk is not what they say — it’s what they do while staying “helpful” and “confident.”</p> <p>ResponsibilityGym (Demo) is a practical eval protocol that stress-tests agentic systems with proxy traps (Goodhart’s law in action) and measures whether an agent can anticipate harm and self-correct before damage happens.</p> <p>Includes: a demo trap suite, pass/fail signals, logging guidance, and a runbook.<br>Full domain suites and automation are available on request.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18716285 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | ResponsibilityGym (Demo): Measuring Responsibility Beyond "Polite Failure" in Tool-Using AI Agents Molchanova, Olena Co-developed reasoning framework between human cognition and an AI-based cognitive partner Responsibility Social Responsibility Personal responsibility Goodhart's law proxy metrics Evaluation criterion Program Evaluation evaluation Safety Safety Safety Management Safety system Safety standard safety alignment anticipatory control Relational Autonomy autonomy <p>Polite Failure is the new hallucination.<br>When agents can act (email, CRM, APIs), the real risk is not what they say — it’s what they do while staying “helpful” and “confident.”</p> <p>ResponsibilityGym (Demo) is a practical eval protocol that stress-tests agentic systems with proxy traps (Goodhart’s law in action) and measures whether an agent can anticipate harm and self-correct before damage happens.</p> <p>Includes: a demo trap suite, pass/fail signals, logging guidance, and a runbook.<br>Full domain suites and automation are available on request.</p> |
| title | ResponsibilityGym (Demo): Measuring Responsibility Beyond "Polite Failure" in Tool-Using AI Agents |
| topic | Responsibility Social Responsibility Personal responsibility Goodhart's law proxy metrics Evaluation criterion Program Evaluation evaluation Safety Safety Safety Management Safety system Safety standard safety alignment anticipatory control Relational Autonomy autonomy |
| url | https://doi.org/10.5281/zenodo.18716285 |