Concrete Problems in AI Safety, Revisited
Fuente:
arXiv
Saved in:
| Main Authors: | Raji, Inioluwa Deborah, Dobbe, Roel |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025)
by: Dobbe, Roel
Published: (2025)
Evaluating Prediction-based Interventions with Human Decision Makers In Mind
by: Raji, Inioluwa Deborah, et al.
Published: (2025)
by: Raji, Inioluwa Deborah, et al.
Published: (2025)
From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems
by: Rismani, Shalaleh, et al.
Published: (2024)
by: Rismani, Shalaleh, et al.
Published: (2024)
AI auditing: The Broken Bus on the Road to AI Accountability
by: Birhane, Abeba, et al.
Published: (2024)
by: Birhane, Abeba, et al.
Published: (2024)
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
by: Ojewale, Victor, et al.
Published: (2024)
by: Ojewale, Victor, et al.
Published: (2024)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
by: Rios-Sialer, Ian
Published: (2026)
by: Rios-Sialer, Ian
Published: (2026)
Open Problems in Machine Unlearning for AI Safety
by: Barez, Fazl, et al.
Published: (2025)
by: Barez, Fazl, et al.
Published: (2025)
Aggregated Individual Reporting for Post-Deployment Evaluation
by: Dai, Jessica, et al.
Published: (2025)
by: Dai, Jessica, et al.
Published: (2025)
From Individual Experience to Collective Evidence: A Reporting-Based Framework for Identifying Systemic Harms
by: Dai, Jessica, et al.
Published: (2025)
by: Dai, Jessica, et al.
Published: (2025)
International Scientific Report on the Safety of Advanced AI (Interim Report)
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
AI Procurement Checklists: Revisiting Implementation in the Age of AI Governance
by: Zick, Tom, et al.
Published: (2024)
by: Zick, Tom, et al.
Published: (2024)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
by: Scholefield, Rebecca, et al.
Published: (2025)
by: Scholefield, Rebecca, et al.
Published: (2025)
Safety Cases: A Scalable Approach to Frontier AI Safety
by: Hilton, Benjamin, et al.
Published: (2025)
by: Hilton, Benjamin, et al.
Published: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
The Data Addition Dilemma
by: Shen, Judy Hanwen, et al.
Published: (2024)
by: Shen, Judy Hanwen, et al.
Published: (2024)
Toward an African Agenda for AI Safety
by: Segun, Samuel T., et al.
Published: (2025)
by: Segun, Samuel T., et al.
Published: (2025)
AI Safety: Necessary, but insufficient and possibly problematic
by: P, Deepak
Published: (2024)
by: P, Deepak
Published: (2024)
Emerging Practices in Frontier AI Safety Frameworks
by: Buhl, Marie Davidsen, et al.
Published: (2025)
by: Buhl, Marie Davidsen, et al.
Published: (2025)
International AI Safety Report
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
by: Costa, Mariana Lins
Published: (2026)
by: Costa, Mariana Lins
Published: (2026)
Upstream and Downstream AI Safety: Both on the Same River?
by: McDermid, John, et al.
Published: (2024)
by: McDermid, John, et al.
Published: (2024)
Combining Cost-Constrained Runtime Monitors for AI Safety
by: Hua, Tim Tian, et al.
Published: (2025)
by: Hua, Tim Tian, et al.
Published: (2025)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
by: Li, Miles Q., et al.
Published: (2026)
by: Li, Miles Q., et al.
Published: (2026)
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Probabilistic Analysis of Copyright Disputes and Generative AI Safety
by: Chiba-Okabe, Hiroaki
Published: (2024)
by: Chiba-Okabe, Hiroaki
Published: (2024)
Building Effective Safety Guardrails in AI Education Tools
by: Clark, Hannah-Beth, et al.
Published: (2025)
by: Clark, Hannah-Beth, et al.
Published: (2025)
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
by: Chin, Yik Chan, et al.
Published: (2026)
by: Chin, Yik Chan, et al.
Published: (2026)
What Is AI Safety? What Do We Want It to Be?
by: Harding, Jacqueline, et al.
Published: (2025)
by: Harding, Jacqueline, et al.
Published: (2025)
The Singapore Consensus on Global AI Safety Research Priorities
by: Bengio, Yoshua, et al.
Published: (2025)
by: Bengio, Yoshua, et al.
Published: (2025)
The Elephant in the Room -- Why AI Safety Demands Diverse Teams
by: Rostcheck, David, et al.
Published: (2024)
by: Rostcheck, David, et al.
Published: (2024)
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
by: Han, Shanshan
Published: (2024)
by: Han, Shanshan
Published: (2024)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
by: Staufer, Leon, et al.
Published: (2026)
by: Staufer, Leon, et al.
Published: (2026)
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
by: Menke, Antonio-Gabriel Chacón, et al.
Published: (2025)
AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations
by: Bogart, Christopher, et al.
Published: (2025)
by: Bogart, Christopher, et al.
Published: (2025)
Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles
by: Mehrotra, Siddharth, et al.
Published: (2025)
by: Mehrotra, Siddharth, et al.
Published: (2025)
Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
by: Reddy, Pavan, et al.
Published: (2025)
by: Reddy, Pavan, et al.
Published: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
by: Brophy, Matthew
Published: (2025)
by: Brophy, Matthew
Published: (2025)
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
by: Petrova, Nora, et al.
Published: (2026)
by: Petrova, Nora, et al.
Published: (2026)
Responsible Data Stewardship: Generative AI and the Digital Waste Problem
by: Utz, Vanessa
Published: (2025)
by: Utz, Vanessa
Published: (2025)
Building Trust: Foundations of Security, Safety and Transparency in AI
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
by: Sidhpurwala, Huzaifa, et al.
Published: (2024)
Similar Items
-
AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
by: Dobbe, Roel
Published: (2025) -
Evaluating Prediction-based Interventions with Human Decision Makers In Mind
by: Raji, Inioluwa Deborah, et al.
Published: (2025) -
From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems
by: Rismani, Shalaleh, et al.
Published: (2024) -
AI auditing: The Broken Bus on the Road to AI Accountability
by: Birhane, Abeba, et al.
Published: (2024) -
Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
by: Ojewale, Victor, et al.
Published: (2024)