Affirmative safety: An approach to risk management for high-risk AI
Fuente:
arXiv
Saved in:
| Main Authors: | Wasil, Akash R., Clymer, Joshua, Krueger, David, Dardaman, Emily, Campos, Simeon, Murphy, Evan R. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
US-China perspectives on extreme AI risks and global governance
by: Wasil, Akash, et al.
Published: (2024)
by: Wasil, Akash, et al.
Published: (2024)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Verification methods for international AI agreements
by: Wasil, Akash R., et al.
Published: (2024)
by: Wasil, Akash R., et al.
Published: (2024)
Combatting deepfakes: Policies to address national security threats and rights violations
by: Miotti, Andrea, et al.
Published: (2024)
by: Miotti, Andrea, et al.
Published: (2024)
Governing dual-use technologies: Case studies of international security agreements and lessons for AI governance
by: Wasil, Akash R., et al.
Published: (2024)
by: Wasil, Akash R., et al.
Published: (2024)
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
by: Coggins, Sam, et al.
Published: (2025)
by: Coggins, Sam, et al.
Published: (2025)
Supervision policies can shape long-term risk management in general-purpose AI models
by: Cebrian, Manuel, et al.
Published: (2025)
by: Cebrian, Manuel, et al.
Published: (2025)
Managing extreme AI risks amid rapid progress
by: Bengio, Yoshua, et al.
Published: (2023)
by: Bengio, Yoshua, et al.
Published: (2023)
Governing frontier general-purpose AI in the public sector: adaptive risk management and policy capacity under uncertainty through 2030
by: Xavier, Fabio Correa
Published: (2026)
by: Xavier, Fabio Correa
Published: (2026)
The global consensus on the risk management of autonomous driving
by: Krügel, Sebastian, et al.
Published: (2025)
by: Krügel, Sebastian, et al.
Published: (2025)
Contemporary AI foundation models increase biological weapons risk
by: Brent, Roger, et al.
Published: (2025)
by: Brent, Roger, et al.
Published: (2025)
What AI evaluations for preventing catastrophic risks can and cannot do
by: Barnett, Peter, et al.
Published: (2024)
by: Barnett, Peter, et al.
Published: (2024)
Episodic memory in AI agents poses risks that should be studied and mitigated
by: DeChant, Chad
Published: (2025)
by: DeChant, Chad
Published: (2025)
Human services organizations and the responsible integration of AI: Considering ethics and contextualizing risk(s)
by: Perron, Brian E., et al.
Published: (2025)
by: Perron, Brian E., et al.
Published: (2025)
It's complicated. The relationship of algorithmic fairness and non-discrimination provisions for high-risk systems in the EU AI Act
by: Meding, Kristof
Published: (2025)
by: Meding, Kristof
Published: (2025)
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
by: Stelling, Lily, et al.
Published: (2025)
by: Stelling, Lily, et al.
Published: (2025)
AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models
by: Barrett, Anthony M., et al.
Published: (2025)
by: Barrett, Anthony M., et al.
Published: (2025)
AI Emergency Preparedness: Examining the federal government's ability to detect and respond to AI-related national security threats
by: Wasil, Akash, et al.
Published: (2024)
by: Wasil, Akash, et al.
Published: (2024)
Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models
by: Barrett, Anthony M., et al.
Published: (2024)
by: Barrett, Anthony M., et al.
Published: (2024)
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
by: Dung, Leonard, et al.
Published: (2025)
by: Dung, Leonard, et al.
Published: (2025)
Visibility into AI Agents
by: Chan, Alan, et al.
Published: (2024)
by: Chan, Alan, et al.
Published: (2024)
Laypeople's Attitudes Towards Fair, Affirmative, and Discriminatory Decision-Making Algorithms
by: Lima, Gabriel, et al.
Published: (2025)
by: Lima, Gabriel, et al.
Published: (2025)
A cross-regional review of AI safety regulations in the commercial aviation
by: Barr, Penny A., et al.
Published: (2025)
by: Barr, Penny A., et al.
Published: (2025)
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
by: Javed, Rafiya, et al.
Published: (2025)
by: Javed, Rafiya, et al.
Published: (2025)
How Large Language Models are Designed to Hallucinate
by: Ackermann, Richard, et al.
Published: (2025)
by: Ackermann, Richard, et al.
Published: (2025)
Recent Advancements In The Field Of Deepfake Detection
by: Krueger, Natalie, et al.
Published: (2023)
by: Krueger, Natalie, et al.
Published: (2023)
Trust in AI: Progress, Challenges, and Future Directions
by: Afroogh, Saleh, et al.
Published: (2024)
by: Afroogh, Saleh, et al.
Published: (2024)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
The AI Criminal Mastermind
by: Krook, Joshua
Published: (2026)
by: Krook, Joshua
Published: (2026)
The potential functions of an international institution for AI safety. Insights from adjacent policy areas and recent trends
by: De Castris, A. Leone, et al.
Published: (2024)
by: De Castris, A. Leone, et al.
Published: (2024)
Subjective Experience in AI Systems: What Do AI Researchers and the Public Believe?
by: Dreksler, Noemi, et al.
Published: (2025)
by: Dreksler, Noemi, et al.
Published: (2025)
Empowering the Future Workforce: Prioritizing Education for the AI-Accelerated Job Market
by: Amini, Lisa, et al.
Published: (2025)
by: Amini, Lisa, et al.
Published: (2025)
Limits of trust in medical AI
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
An investigation of AI integration in sound designer workflows and experiences
by: Garcia, Nelly, et al.
Published: (2026)
by: Garcia, Nelly, et al.
Published: (2026)
When Autonomy Breaks: The Hidden Existential Risk of AI
by: Krook, Joshua
Published: (2025)
by: Krook, Joshua
Published: (2025)
Reproducible workflow for online AI in digital health
by: Ghosh, Susobhan, et al.
Published: (2025)
by: Ghosh, Susobhan, et al.
Published: (2025)
AI Agents and the Law
by: Riedl, Mark O., et al.
Published: (2025)
by: Riedl, Mark O., et al.
Published: (2025)
Beyond Agreement: Rethinking Ground Truth in Educational AI Annotation
by: Thomas, Danielle R., et al.
Published: (2025)
by: Thomas, Danielle R., et al.
Published: (2025)
Manipulation and the AI Act: Large Language Model Chatbots and the Danger of Mirrors
by: Krook, Joshua
Published: (2025)
by: Krook, Joshua
Published: (2025)
Similar Items
-
US-China perspectives on extreme AI risks and global governance
by: Wasil, Akash, et al.
Published: (2024) -
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024) -
Verification methods for international AI agreements
by: Wasil, Akash R., et al.
Published: (2024) -
Combatting deepfakes: Policies to address national security threats and rights violations
by: Miotti, Andrea, et al.
Published: (2024) -
Governing dual-use technologies: Case studies of international security agreements and lessons for AI governance
by: Wasil, Akash R., et al.
Published: (2024)