To Err is AI : A Case Study Informing LLM Flaw Reporting Practices
Fuente:
arXiv
Guardado en:
| Autores principales: | McGregor, Sean, Ettinger, Allyson, Judd, Nick, Albee, Paul, Jiang, Liwei, Rao, Kavel, Smith, Will, Longpre, Shayne, Ghosh, Avijit, Fiorelli, Christopher, Hoang, Michelle, Cattell, Sven, Dziri, Nouha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
por: Han, Seungju, et al.
Publicado: (2024)
por: Han, Seungju, et al.
Publicado: (2024)
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
por: Cattell, Sven, et al.
Publicado: (2024)
por: Cattell, Sven, et al.
Publicado: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
por: Jiang, Liwei, et al.
Publicado: (2024)
por: Jiang, Liwei, et al.
Publicado: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
por: Rao, Kavel, et al.
Publicado: (2023)
por: Rao, Kavel, et al.
Publicado: (2023)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
por: Rabin, Rafiqul, et al.
Publicado: (2025)
por: Rabin, Rafiqul, et al.
Publicado: (2025)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
por: Oderinwale, Hamidah, et al.
Publicado: (2024)
por: Oderinwale, Hamidah, et al.
Publicado: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
por: Lu, Ximing, et al.
Publicado: (2024)
por: Lu, Ximing, et al.
Publicado: (2024)
Surfacing Semantic Orthogonality Across Model Safety Benchmarks: A Multi-Dimensional Analysis
por: Bennion, Jonathan, et al.
Publicado: (2025)
por: Bennion, Jonathan, et al.
Publicado: (2025)
SandboxEval: Towards Securing Test Environment for Untrusted Code
por: Rabin, Rafiqul, et al.
Publicado: (2025)
por: Rabin, Rafiqul, et al.
Publicado: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
por: Sorensen, Taylor, et al.
Publicado: (2023)
por: Sorensen, Taylor, et al.
Publicado: (2023)
In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI
por: Longpre, Shayne, et al.
Publicado: (2025)
por: Longpre, Shayne, et al.
Publicado: (2025)
A Systematic Review of NeurIPS Dataset Management Practices
por: Wu, Yiwei, et al.
Publicado: (2024)
por: Wu, Yiwei, et al.
Publicado: (2024)
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
por: Longpre, Shayne, et al.
Publicado: (2025)
por: Longpre, Shayne, et al.
Publicado: (2025)
Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
por: Misra, Kanishka, et al.
Publicado: (2024)
por: Misra, Kanishka, et al.
Publicado: (2024)
When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
por: Li, Yanhong, et al.
Publicado: (2024)
por: Li, Yanhong, et al.
Publicado: (2024)
La historia del zoológico : Caja : Citas del presidene Mao Zedong / Edward Albee ; traducción y prólogo Víctor Weinstock ; introducción Hern n Lara Zavala
por: Albee, Edward
por: Albee, Edward
Do environmental shocks create new coalitions? The development of a contingent coalition after Three Mile Island
por: Jasper Cattell
Publicado: (2025)
por: Jasper Cattell
Publicado: (2025)
AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research
por: Simmons-Edler, Riley, et al.
Publicado: (2024)
por: Simmons-Edler, Riley, et al.
Publicado: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
por: Graf, Victoria, et al.
Publicado: (2026)
por: Graf, Victoria, et al.
Publicado: (2026)
ColorGrid: A Multi-Agent Non-Stationary Environment for Goal Inference and Assistance
por: Risukhin, Andrey, et al.
Publicado: (2025)
por: Risukhin, Andrey, et al.
Publicado: (2025)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
por: Zhou, Kaitlyn, et al.
Publicado: (2024)
por: Zhou, Kaitlyn, et al.
Publicado: (2024)
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
por: Jiang, Liwei, et al.
Publicado: (2025)
por: Jiang, Liwei, et al.
Publicado: (2025)
Expediency-Based Practice? Medical Students' Reliance on Google and Wikipedia for Biomedical Inquiries
por: Judd, Terry, et al.
Publicado: (2011)
por: Judd, Terry, et al.
Publicado: (2011)
Flexible Audit Trailing in Interactive Courseware.
por: Judd, Terry, et al.
Publicado: (2001)
por: Judd, Terry, et al.
Publicado: (2001)
Tradition und Diskurs
por: Dziri, Amir
Publicado: (2023)
por: Dziri, Amir
Publicado: (2023)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
por: Sun, Yiyou, et al.
Publicado: (2025)
por: Sun, Yiyou, et al.
Publicado: (2025)
North Pacific Ocean gyre primary production (PP) from 1969-70 based on 14C uptake
por: Cattell, S A, et al.
Publicado: (2003)
por: Cattell, S A, et al.
Publicado: (2003)
To Err Is Human, but Llamas Can Learn It Too
por: Luhtaru, Agnes, et al.
Publicado: (2024)
por: Luhtaru, Agnes, et al.
Publicado: (2024)
CULTURE-GEN: Revealing Global Cultural Perception in Language Models through Natural Language Prompting
por: Li, Huihan, et al.
Publicado: (2024)
por: Li, Huihan, et al.
Publicado: (2024)
Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
por: Qiu, Linlu, et al.
Publicado: (2023)
por: Qiu, Linlu, et al.
Publicado: (2023)
The 2024 Foundation Model Transparency Index
por: Bommasani, Rishi, et al.
Publicado: (2024)
por: Bommasani, Rishi, et al.
Publicado: (2024)
Quantum fluctuation dynamics of open quantum systems with collective operator-valued rates, and applications to Hopfield-like networks
por: Fiorelli, Eliana
Publicado: (2024)
por: Fiorelli, Eliana
Publicado: (2024)
PROPRIEDADES MECÂNICAS DE PEÇAS COM DIMENSÕES ESTRUTURAIS DE Pinus spp: CORRELAÇÃO ENTRE RESISTÊNCIA À TRAÇÃO E CLASSIFICAÇÃO VISUAL
por: Juliano Fiorelli
Publicado: (2009)
por: Juliano Fiorelli
Publicado: (2009)
Particleboards with waste wood from reforestation
por: Juliano Fiorelli
Publicado: (2014)
por: Juliano Fiorelli
Publicado: (2014)
Painéis de partículas à base de bagaço de cana e resina de mamona - produção e propriedades
por: Juliano Fiorelli
Publicado: (2011)
por: Juliano Fiorelli
Publicado: (2011)
AI for Scientific Discovery is a Social Problem
por: Channing, Georgia, et al.
Publicado: (2025)
por: Channing, Georgia, et al.
Publicado: (2025)
To Err is Machine: Vulnerability Detection Challenges LLM Reasoning
por: Steenhoek, Benjamin, et al.
Publicado: (2024)
por: Steenhoek, Benjamin, et al.
Publicado: (2024)
Err on the Side of Texture: Texture Bias on Real Data
por: Hoak, Blaine, et al.
Publicado: (2024)
por: Hoak, Blaine, et al.
Publicado: (2024)
Ejemplares similares
-
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
por: Han, Seungju, et al.
Publicado: (2024) -
Coordinated Flaw Disclosure for AI: Beyond Security Vulnerabilities
por: Cattell, Sven, et al.
Publicado: (2024) -
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
por: Jiang, Liwei, et al.
Publicado: (2024) -
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
por: Rao, Kavel, et al.
Publicado: (2023) -
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
por: Rabin, Rafiqul, et al.
Publicado: (2025)