AI Safety is Stuck in Technical Terms -- A System Safety Response to the International AI Safety Report
Fuente:
arXiv
Salvato in:
| Autore principale: | Dobbe, Roel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Concrete Problems in AI Safety, Revisited
di: Raji, Inioluwa Deborah, et al.
Pubblicazione: (2023)
di: Raji, Inioluwa Deborah, et al.
Pubblicazione: (2023)
International AI Safety Report
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
International Scientific Report on the Safety of Advanced AI (Interim Report)
di: Bengio, Yoshua, et al.
Pubblicazione: (2024)
di: Bengio, Yoshua, et al.
Pubblicazione: (2024)
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025)
Safety Cases: How to Justify the Safety of Advanced AI Systems
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
di: Staufer, Leon, et al.
Pubblicazione: (2026)
di: Staufer, Leon, et al.
Pubblicazione: (2026)
Safety Cases: A Scalable Approach to Frontier AI Safety
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
di: Hilton, Benjamin, et al.
Pubblicazione: (2025)
Human-AI Safety: A Descendant of Generative AI and Control Systems Safety
di: Bajcsy, Andrea, et al.
Pubblicazione: (2024)
di: Bajcsy, Andrea, et al.
Pubblicazione: (2024)
Toward an African Agenda for AI Safety
di: Segun, Samuel T., et al.
Pubblicazione: (2025)
di: Segun, Samuel T., et al.
Pubblicazione: (2025)
Questionnaire Responses Do not Capture the Safety of AI Agents
di: Hellrigel-Holderbaum, Max, et al.
Pubblicazione: (2026)
di: Hellrigel-Holderbaum, Max, et al.
Pubblicazione: (2026)
Emerging Practices in Frontier AI Safety Frameworks
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
di: Buhl, Marie Davidsen, et al.
Pubblicazione: (2025)
AI Safety: Necessary, but insufficient and possibly problematic
di: P, Deepak
Pubblicazione: (2024)
di: P, Deepak
Pubblicazione: (2024)
Agentic Microphysics: A Manifesto for Generative AI Safety
di: Pierucci, Federico, et al.
Pubblicazione: (2026)
di: Pierucci, Federico, et al.
Pubblicazione: (2026)
From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems
di: Rismani, Shalaleh, et al.
Pubblicazione: (2024)
di: Rismani, Shalaleh, et al.
Pubblicazione: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
di: Ren, Richard, et al.
Pubblicazione: (2024)
di: Ren, Richard, et al.
Pubblicazione: (2024)
Combining Cost-Constrained Runtime Monitors for AI Safety
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
Building Effective Safety Guardrails in AI Education Tools
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
di: Clark, Hannah-Beth, et al.
Pubblicazione: (2025)
What Is AI Safety? What Do We Want It to Be?
di: Harding, Jacqueline, et al.
Pubblicazione: (2025)
di: Harding, Jacqueline, et al.
Pubblicazione: (2025)
The Singapore Consensus on Global AI Safety Research Priorities
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
di: Bengio, Yoshua, et al.
Pubblicazione: (2025)
The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations
di: Costa, Mariana Lins
Pubblicazione: (2026)
di: Costa, Mariana Lins
Pubblicazione: (2026)
Upstream and Downstream AI Safety: Both on the Same River?
di: McDermid, John, et al.
Pubblicazione: (2024)
di: McDermid, John, et al.
Pubblicazione: (2024)
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
di: Li, Miles Q., et al.
Pubblicazione: (2026)
di: Li, Miles Q., et al.
Pubblicazione: (2026)
Probabilistic Analysis of Copyright Disputes and Generative AI Safety
di: Chiba-Okabe, Hiroaki
Pubblicazione: (2024)
di: Chiba-Okabe, Hiroaki
Pubblicazione: (2024)
Interoperability in AI Safety Governance: Ethics, Regulations, and Standards
di: Chin, Yik Chan, et al.
Pubblicazione: (2026)
di: Chin, Yik Chan, et al.
Pubblicazione: (2026)
Introduction to AI Safety, Ethics, and Society
di: Hendrycks, Dan
Pubblicazione: (2024)
di: Hendrycks, Dan
Pubblicazione: (2024)
AI Safety as Control of Irreversibility: A Systems Framework for Decision-Energy and Sovereignty Boundaries
di: Shu, Wesley, et al.
Pubblicazione: (2026)
di: Shu, Wesley, et al.
Pubblicazione: (2026)
An Approach to Technical AGI Safety and Security
di: Shah, Rohin, et al.
Pubblicazione: (2025)
di: Shah, Rohin, et al.
Pubblicazione: (2025)
AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
di: Bowen, Dillon, et al.
Pubblicazione: (2025)
di: Bowen, Dillon, et al.
Pubblicazione: (2025)
The Elephant in the Room -- Why AI Safety Demands Diverse Teams
di: Rostcheck, David, et al.
Pubblicazione: (2024)
di: Rostcheck, David, et al.
Pubblicazione: (2024)
Bridging Today and the Future of Humanity: AI Safety in 2024 and Beyond
di: Han, Shanshan
Pubblicazione: (2024)
di: Han, Shanshan
Pubblicazione: (2024)
Annotating the Chain-of-Thought: A Behavior-Labeled Dataset for AI Safety
di: Menke, Antonio-Gabriel Chacón, et al.
Pubblicazione: (2025)
di: Menke, Antonio-Gabriel Chacón, et al.
Pubblicazione: (2025)
Building Trust: Foundations of Security, Safety and Transparency in AI
di: Sidhpurwala, Huzaifa, et al.
Pubblicazione: (2024)
di: Sidhpurwala, Huzaifa, et al.
Pubblicazione: (2024)
AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
di: Lin, Zhiqiang, et al.
Pubblicazione: (2025)
di: Lin, Zhiqiang, et al.
Pubblicazione: (2025)
Open Problems in Machine Unlearning for AI Safety
di: Barez, Fazl, et al.
Pubblicazione: (2025)
di: Barez, Fazl, et al.
Pubblicazione: (2025)
The Homogenization Problem in LLMs: Towards Meaningful Diversity in AI Safety
di: Rios-Sialer, Ian
Pubblicazione: (2026)
di: Rios-Sialer, Ian
Pubblicazione: (2026)
Preventing Another Tessa: Modular Safety Middleware For Health-Adjacent AI Assistants
di: Reddy, Pavan, et al.
Pubblicazione: (2025)
di: Reddy, Pavan, et al.
Pubblicazione: (2025)
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
di: Brophy, Matthew
Pubblicazione: (2025)
di: Brophy, Matthew
Pubblicazione: (2025)
The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries
di: Petrova, Nora, et al.
Pubblicazione: (2026)
di: Petrova, Nora, et al.
Pubblicazione: (2026)
International AI Safety Report 2026
di: Bengio, Yoshua, et al.
Pubblicazione: (2026)
di: Bengio, Yoshua, et al.
Pubblicazione: (2026)
AI Safety Should Prioritize the Future of Work
di: Hazra, Sanchaita, et al.
Pubblicazione: (2025)
di: Hazra, Sanchaita, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Concrete Problems in AI Safety, Revisited
di: Raji, Inioluwa Deborah, et al.
Pubblicazione: (2023) -
International AI Safety Report
di: Bengio, Yoshua, et al.
Pubblicazione: (2025) -
International Scientific Report on the Safety of Advanced AI (Interim Report)
di: Bengio, Yoshua, et al.
Pubblicazione: (2024) -
International Agreements on AI Safety: Review and Recommendations for a Conditional AI Safety Treaty
di: Scholefield, Rebecca, et al.
Pubblicazione: (2025) -
Safety Cases: How to Justify the Safety of Advanced AI Systems
di: Clymer, Joshua, et al.
Pubblicazione: (2024)