To Err is AI : A Case Study Informing LLM Flaw Reporting Practices
Fuente:
arXiv
Saved in:
| Main Authors: | McGregor, Sean, Ettinger, Allyson, Judd, Nick, Albee, Paul, Jiang, Liwei, Rao, Kavel, Smith, Will, Longpre, Shayne, Ghosh, Avijit, Fiorelli, Christopher, Hoang, Michelle, Cattell, Sven, Dziri, Nouha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024)
by: Han, Seungju, et al.
Published: (2024)
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
by: Cattell, Sven, et al.
Published: (2024)
by: Cattell, Sven, et al.
Published: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
by: Jiang, Liwei, et al.
Published: (2024)
by: Jiang, Liwei, et al.
Published: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
by: Rao, Kavel, et al.
Published: (2023)
by: Rao, Kavel, et al.
Published: (2023)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
by: Rabin, Rafiqul, et al.
Published: (2025)
by: Rabin, Rafiqul, et al.
Published: (2025)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
by: Oderinwale, Hamidah, et al.
Published: (2024)
by: Oderinwale, Hamidah, et al.
Published: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
by: Lu, Ximing, et al.
Published: (2024)
by: Lu, Ximing, et al.
Published: (2024)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
by: Bennion, Jonathan, et al.
Published: (2025)
by: Bennion, Jonathan, et al.
Published: (2025)
SandboxEval: Towards Securing Test Environment for Untrusted Code
by: Rabin, Rafiqul, et al.
Published: (2025)
by: Rabin, Rafiqul, et al.
Published: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
by: Sorensen, Taylor, et al.
Published: (2023)
by: Sorensen, Taylor, et al.
Published: (2023)
In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
A Systematic Review of NeurIPS Dataset Management Practices
by: Wu, Yiwei, et al.
Published: (2024)
by: Wu, Yiwei, et al.
Published: (2024)
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
by: Longpre, Shayne, et al.
Published: (2025)
by: Longpre, Shayne, et al.
Published: (2025)
Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
by: Misra, Kanishka, et al.
Published: (2024)
by: Misra, Kanishka, et al.
Published: (2024)
When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
by: Li, Yanhong, et al.
Published: (2024)
by: Li, Yanhong, et al.
Published: (2024)
La historia del zoológico : Caja : Citas del presidene Mao Zedong / Edward Albee ; traducción y prólogo Víctor Weinstock ; introducción Hern n Lara Zavala
by: Albee, Edward
by: Albee, Edward
Do environmental shocks create new coalitions? The development of a contingent coalition after Three Mile Island
by: Jasper Cattell
Published: (2025)
by: Jasper Cattell
Published: (2025)
AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research
by: Simmons-Edler, Riley, et al.
Published: (2024)
by: Simmons-Edler, Riley, et al.
Published: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
by: Graf, Victoria, et al.
Published: (2026)
by: Graf, Victoria, et al.
Published: (2026)
ColorGrid: A Multi-Agent Non-Stationary Environment for Goal Inference and Assistance
by: Risukhin, Andrey, et al.
Published: (2025)
by: Risukhin, Andrey, et al.
Published: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
by: Zhou, Kaitlyn, et al.
Published: (2024)
by: Zhou, Kaitlyn, et al.
Published: (2024)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
by: Jiang, Liwei, et al.
Published: (2025)
by: Jiang, Liwei, et al.
Published: (2025)
Expediency-Based Practice? Medical Students' Reliance on Google and Wikipedia for Biomedical Inquiries
by: Judd, Terry, et al.
Published: (2011)
by: Judd, Terry, et al.
Published: (2011)
Flexible Audit Trailing in Interactive Courseware.
by: Judd, Terry, et al.
Published: (2001)
by: Judd, Terry, et al.
Published: (2001)
Tradition und Diskurs
by: Dziri, Amir
Published: (2023)
by: Dziri, Amir
Published: (2023)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
North Pacific Ocean gyre primary production (PP) from 1969-70 based on 14C uptake
by: Cattell, S A, et al.
Published: (2003)
by: Cattell, S A, et al.
Published: (2003)
To Err Is Human, but Llamas Can Learn It Too
by: Luhtaru, Agnes, et al.
Published: (2024)
by: Luhtaru, Agnes, et al.
Published: (2024)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
by: Li, Huihan, et al.
Published: (2024)
by: Li, Huihan, et al.
Published: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
by: Qiu, Linlu, et al.
Published: (2023)
by: Qiu, Linlu, et al.
Published: (2023)
The 2024 Foundation Model Transparency Index
by: Bommasani, Rishi, et al.
Published: (2024)
by: Bommasani, Rishi, et al.
Published: (2024)
Quantum fluctuation dynamics of open quantum systems with collective operator-valued rates, and applications to Hopfield-like networks
by: Fiorelli, Eliana
Published: (2024)
by: Fiorelli, Eliana
Published: (2024)
PROPRIEDADES MECÂNICAS DE PEÇAS COM DIMENSÕES ESTRUTURAIS DE Pinus spp: CORRELAÇÃO ENTRE RESISTÊNCIA À TRAÇÃO E CLASSIFICAÇÃO VISUAL
by: Juliano Fiorelli
Published: (2009)
by: Juliano Fiorelli
Published: (2009)
Particleboards with waste wood from reforestation
by: Juliano Fiorelli
Published: (2014)
by: Juliano Fiorelli
Published: (2014)
Painéis de partículas à base de bagaço de cana e resina de mamona - produção e propriedades
by: Juliano Fiorelli
Published: (2011)
by: Juliano Fiorelli
Published: (2011)
AI for Scientific Discovery is a Social Problem
by: Channing, Georgia, et al.
Published: (2025)
by: Channing, Georgia, et al.
Published: (2025)
To Err is Machine: Vulnerability Detection Challenges LLM Reasoning
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
Err on the Side of Texture: Texture Bias on Real Data
by: Hoak, Blaine, et al.
Published: (2024)
by: Hoak, Blaine, et al.
Published: (2024)
Similar Items
-
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
by: Han, Seungju, et al.
Published: (2024) -
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
by: Cattell, Sven, et al.
Published: (2024) -
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
by: Jiang, Liwei, et al.
Published: (2024) -
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
by: Rao, Kavel, et al.
Published: (2023) -
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
by: Rabin, Rafiqul, et al.
Published: (2025)